Site Reliability Engineer
ArangoDB
WorldwideremotePosted 1 month ago
Skill Required
Site-Reliability-EngineeringDevOps-EngineerCloud-Infrastructure-EngineerPlatform-EngineerSRE-EngineerSenior-Site-Reliability-EngineerPrincipal-Site-Reliability-EngineerSite-Reliability-Engineering-LeadSite Reliability EngineerSystem DesignBackup and RecoveryDockerObservabilityCD pipelinesGCPR ProgrammingKubernetesPrometheusnetworkingTerraformsecuritybuildingGitHub ActionsJenkinsetc.)PythonGolangdesignDevOpsArgoCDLinuxCI/CDCloudDesign PatternsFulltime
Key highlights
- Remote-first position in EU Timezone
- Focus on Kubernetes and cloud-native infrastructure
- Requires Golang proficiency or willingness to learn
- Collaborative and growth-oriented company culture
Role overview
ArangoDB is seeking a Site Reliability Engineer (SRE) to ensure the reliability, scalability, and performance of their robust, cloud-native infrastructure. The role focuses on automating, monitoring, and optimizing distributed database systems that power mission-critical applications across various industries.
Responsibilities
- Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms.
- Ensure the scalability, performance, and reliability of our Kubernetes-based distributed database systems.
- Collaborate with developers to write efficient, production-grade code in Golang to automate infrastructure management and improve system operations.
- Optimize and automate CI/CD pipelines, deployment processes, and monitoring systems to support our production environment.
- Develop strategies for disaster recovery, high availability, and fault tolerance.
- Proactively identify system bottlenecks, troubleshoot, and resolve issues across the stack (network, OS, cloud infrastructure).
- Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance.
- Participate in on-call rotations to support critical production systems and respond to incidents.
- Collaborate with cross-functional teams to improve overall system reliability and scalability.
- Collaborate with the Customer Success team to resolve customer issues.
Requirements
- Proven experience as an SRE or DevOps Engineer in a cloud-native environment.
- Proficiency with Kubernetes in managing large-scale, distributed systems.
- Experience with cloud providers such as AWS and Google Cloud (GCP).
- Solid understanding of networking, security practices, and troubleshooting methods.
- Understanding of Linux internals (processes, environment variables etc.).
- Familiarity with containerization technologies (e.g., Docker).
- Knowledge of CI/CD practices and tools (Jenkins, CircleCI, etc.).
- Familiarity with alerting, monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack).
- Strong troubleshooting and problem-solving skills, with the ability to address complex infrastructure issues.
- Excellent communication and collaboration skills with a focus on continuous improvement and operational excellence.
- Strong ability to self-organize and to work independently as part of a remote team.
- Knowledge of version control systems, particularly Git.
- Familiarity with programming languages such as Golang or Python.
Nice to have
- Experience managing distributed databases or large-scale data storage systems.
- Knowledge of security best practices in cloud environments.
- Experience with scripting languages like Python or Bash.
- Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus.
- Experience working with GitOps.
- Strong programming skills in Golang, with experience in developing automation tools, scripts, or services.
Benefits
- Supportive work environment where employees can grow, learn, and share knowledge.
- Inclusive, growth-oriented team.
- Opportunity to contribute to cutting-edge AI and data infrastructure.
- Collaboration with experienced engineers, marketers, and product leaders.
- Helping shape how enterprises build AI-powered applications.
Additional details
- Values: Innovation, Customer-Centric Focus, Collaboration and Growth.
- Location: EU Timezone, preferably within the EU itself (Remote).
- Originally posted on Himalayas.