We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to drive the reliability, scalability, security, observability, and performance of mission‑critical production systems. The ideal candidate should have strong expertise in Azure Cloud, Kubernetes, DevOps, SRE practices, and modern observability tools, with the ability to balance production support and long‑term reliability engineering initiatives.
Responsibilities
- Participate in 24/7 production on‑call rotation.
- Troubleshoot and resolve high‑severity production incidents.
- Lead Root Cause Analysis (RCA) and post‑mortem activities.
- Reduce MTTR through proactive reliability improvements.
- Maintain SLAs, SLOs, and service reliability.
- Implement SRE best practices across production environments.
- Define and improve SLIs, SLOs, SLAs, and Error Budgets.
- Reduce operational toil through automation.
- Improve service availability, scalability, and disaster recovery readiness.
- Perform reliability reviews and capacity planning.
- Design, deploy, and manage Azure infrastructure.
- Administer Kubernetes clusters and containerized workloads.
- Manage Infrastructure as Code using Terraform.
- Deploy applications using Helm and Git‑based CI/CD workflows.
- Build and maintain observability solutions using OpenTelemetry.
- Implement monitoring with Prometheus, Grafana, Azure Monitor, and Datadog.
- Establish monitoring based on Golden Signals.
- Improve logging, tracing, alerting, and performance monitoring across distributed systems.
- Implement cloud security best practices.
- Manage IAM, network security, and secrets management.
- Support vulnerability remediation and compliance initiatives.
Requirements
- 10+ years of overall IT experience.
- Minimum 6–7 years of hands‑on DevOps/SRE experience.
- Experience managing large‑scale production environments.
- Immediate joiners preferred.
- Official notice period must be Immediate (already serving notice) or up to 30 days only.
- Strong Azure Cloud experience (Mandatory).
- Microsoft Azure (Mandatory).
- Kubernetes (Strong hands‑on experience).
- Terraform (Infrastructure as Code).
- Helm.
- GitHub / GitLab / Azure Repos.
- Site Reliability Engineering (SRE) Practices.
- Incident Response & 24/7 On‑Call Support.
- Root Cause Analysis (RCA).
- SLI / SLO / SLA Management.
- Error Budgets.
- Capacity Planning.
- Toil Reduction.
- Reliability Engineering.
- Observability & Monitoring.
- OpenTelemetry.
- Prometheus.
- Grafana.
- Azure Monitor.
- Datadog.
- Distributed Tracing.
- Metrics, Logs & Traces.
- Golden Signals Monitoring: Latency, Traffic, Errors, Saturation.
- Python.
- Bash Scripting.
- Linux Administration.
- Networking Fundamentals: DNS, TCP/IP, Load Balancing, SSL/TLS.
- Strong ownership and accountability.
- Excellent debugging and analytical skills.
- Calm decision‑making during production incidents.
- Deep understanding of SRE principles and observability.
- Passion for automation, scalability, and continuous improvement.
- Strong communication and collaboration skills.
Nice to have
- Google Cloud Platform (GCP).
- AWS (EC2, S3, RDS, IAM, VPC, CloudWatch).
- Go (Golang).
- OpenSearch / ELK Stack.
- Azure AI Services.
- AI Foundry.
- AI/ML Infrastructure.
- RAG (Retrieval‑Augmented Generation) Workloads.
- Experience with Distributed Systems Architecture.
Additional details
- Hiring: Senior Site Reliability Engineer (SRE) / DevOps Engineer.
- Location: Viman Nagar, Pune (Work From Office).
- Experience: 10+ Years.
- CTC: Up to ₹25‑27 LPA.
- Joining: Immediate Joiners Preferred.
- Shift Timings: 3:00 PM – 12:00 AM (Monday – Friday).
- On‑Call: 24/7 Production Support (Rotation).
- ⚠️ Strict Screening Note: Candidates must have 8+ years of relevant experience (Overall in IT 10yrs+) with the required mandatory skills. Official notice period must be Immediate or up to 30 days only. Candidates with notice periods exceeding 30 days will not be considered.
- Screening Notes (Mandatory): Immediate joiners only. Minimum 10+ years of experience. Strong Azure Cloud experience (Mandatory). GCP exposure is acceptable as an added advantage. Minimum 6–7 years of hands‑on SRE(Primary)/DevOps experience. Strong expertise in Kubernetes, OpenTelemetry, Golden Signals, Prometheus, Grafana, Terraform, and SRE Practices is mandatory. Candidates should have experience supporting high‑availability production environments and be comfortable with 24/7 on‑call rotation.
- Skills: grafana, prometheus, sre practices, python, sre, opentelemetry, kubernetes, terraform, golden signals, 24/7 on‑call rotation, azure cloud.