Posted today · be early
Site Reliability Engineer - India
JumpCloud
IndiaremotePosted 1 day ago
Skill Required
Site-Reliability-EngineerSRE-EngineerDevOps-EngineerPlatform-EngineerSoftware-EngineerSite-Reliability-Engineering-JobsSite-Reliability-Engineering-(SRE)Site-Reliability-EngineeringSite Reliability Engineersoftware engineeringBackup and RecoverySOCVaultGitHub ActionstroubleshootingMicroservicesObservabilityEngineeringKubernetesnetworkingautomationTerraformdesigningsimilar)buildingDatadogNginxTestNGPythondesignDevOpsArgoCDGitIstioFulltime
Key highlights
- 5+ years of professional software engineering experience
- On-call rotation expectation
- Remote work within specified country
- Must be authorized to work in the country noted in the job description
- English fluency required
- AI-assisted development tools used (Cursor, Claude Code, GitHub Copilot)
Role overview
JumpCloud is seeking a Software Engineer 3 (SRE) to join its Infrastructure & Reliability Engineering team. This role focuses on ensuring high availability, performance, recoverability, and operational maturity across JumpCloud's production systems. The SRE will design automation, build observability frameworks, define reliability standards, and reduce operational toil through code. The position involves building and scaling cloud-native infrastructure, participating in incident management, and implementing reliability best practices across the directory platform and microservices. This is a remote role based in the country specified in the job description, with an expectation to participate in on-call rotations.
Responsibilities
- Design, deploy, and maintain the reliability, availability, and performance of critical JumpCloud systems and APIs across AWS and GCP.
- Operationalize SLIs, SLOs, and error budgets in direct partnership with core application teams.
- Build and refine end-to-end observability across microservices and cloud infrastructure using tools like Datadog.
- Implement actionable monitoring across Golden Signals (Latency, Traffic, Errors, Saturation) to optimize detection (MTTD) and minimize alert fatigue.
- Participate in on-call rotations, incident response, and blameless post-incident reviews to drive continuous systemic improvements.
- Manage and operationalize production Kubernetes (EKS) clusters utilizing GitOps delivery workflows (Argo CD, Kargo).
- Provision and secure multi-cloud infrastructure using modular Terraform (Infrastructure-as-Code).
- Develop and maintain Disaster Recovery (DR) dashboards, runbooks, multi-region failover automation, and validation tests to ensure alignment with defined RTO and RPO targets.
- Eliminate operational toil by writing production-grade Python or Go scripts and automation tools.
- Leverage AI-assisted development tools (Cursor, Claude Code, GitHub Copilot) to accelerate scripting, runbook generation, and incident triage.
Requirements
- 5+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical systems.
- Python/Go Proficiency: Hands-on capabilities writing code for SRE tools, custom automation, and cloud integrations.
- Kubernetes Ecosystem: Production experience with Kubernetes cluster operations, container orchestration, and GitOps pipelines (Argo CD).
- Infrastructure as Code: Solid experience writing, maintaining, and modularizing Terraform configurations.
- Cloud Architecture: Direct experience operating cloud workloads on AWS (EKS, IAM, VPC networking, Route53, ALB/NLB) or GCP.
- FinOps & Cost Visibility: Practical experience setting up cost-allocation tagging, resource right-sizing, and building FinOps dashboards to visualize cloud spend.
- Disaster Recovery & Monitoring: Experience building DR dashboards, running failover drills, and configuring monitoring tools to track system health and recovery metrics.
- Observability & Incident Management: Practical experience with Datadog (or similar), PagerDuty, alerting hygiene, and working within SLI/SLO frameworks.
- Solid operational experience configuring and troubleshooting production service meshes (Istio or similar) and managing high-availability proxy solutions (HAProxy, NGINX, or similar).
- Problem Solving & Mindset: Strong troubleshooting skills, effective collaboration, and a track record of driving operational efficiency through code.
- A strong team player who helps us live by our core values: building connections, thinking big, and getting 1% better every day.
- Must be located in and authorized to work in the country noted in the job description.
- English fluency required for speaking and writing.
Nice to have
- Experience with CI/CD tools such as GitHub Actions or GitLab Pipelines.
- Basic understanding of chaos engineering principles or testing resilience in staging/production.
- Familiarity with secrets management tools (HashiCorp Vault, AWS Secrets Manager, External Secrets Operator).
- Basic knowledge of DevSecOps tools and scanning/fixing infrastructure-as-code vulnerabilities.
Additional details
- All roles at JumpCloud are Remote unless otherwise specified in the Job Description.
- There is an expectation that engineers participate in on-call rotations and be ready and able to respond during assigned shifts.
- JumpCloud has teams in 15+ countries around the world and conducts internal business in English.
- The interview and any additional screening process will take place primarily in English.
- JumpCloud is an equal opportunity employer.
- JumpCloud is not accepting third party resumes at this time.
- Scam Notice: JumpCloud will never ask for personal account information during recruitment, never send checks for equipment, and all official communication comes from @jumpcloud.com addresses.
- To apply, please submit your résumé and brief explanation about yourself and why you would be a good fit for JumpCloud.