Site Reliability Engineer - SRE
Drivetrain
IndiaremotePosted 1 month ago
Skill Required
Site-Reliability-EngineeringSREDevOpsCloud-Infrastructure-EngineeringCloud-EngineerSenior-Site-Reliability-EngineerSite-Reliability-Engineering-ManagerSite-Reliability-Engineering-LeadStaff-Site-Reliability-Engineer-(SRE)Site Reliability EngineerSOCDockerNginxObservabilityCD pipelinesIstioEngineeringR ProgrammingKubernetesPrometheusnetworkingautomationTerraformsecuritybuildingEmbedded CJenkinsTestNGPythondesignCI/CDCloudDesign PatternsRustFulltime
Key highlights
- Remote-first work environment
- 5+ years of SRE/DevOps/Cloud Infrastructure experience required
- Multi-cloud (AWS & GCP) expertise mandatory
- Strong focus on automation, observability, and reliability
Role overview
Drivetrain is on a mission to empower businesses to make better decisions through its financial planning & decision-making platform, helping companies scale and achieve targets predictably. As a Senior Site Reliability Engineer, you will be a cornerstone of the engineering organization, ensuring the fast-growing SaaS platform remains highly available, performant, and secure. You will own multi-cloud infrastructure, drive automation, champion observability, and collaborate with development teams to build a culture of reliability from code commit to production.
Responsibilities
- Architect, manage, and continuously optimize highly available cloud infrastructure across both AWS and GCP, balancing workload demands to ensure maximum cost-efficiency, scalability, and strict security compliance across both platforms.
- Lead the design, deployment, and management of scalable Kubernetes clusters.
- Utilize configuration management tools like Kustomize to enforce standardized, repeatable, and automated deployment configurations across all environments.
- Implement and maintain service mesh technologies (e.g., Istio, Linkerd) to secure, control, and observe service-to-service communication.
- Drive container security best practices, including image scanning, runtime protection, and strict RBAC enforcement.
- Architect, maintain, and optimize robust CI/CD pipelines using Git and Jenkins, focusing on reducing deployment friction, accelerating release velocity, and enforcing automated testing and security gates.
- Write, review, and maintain Terraform modules to provision and manage cloud resources predictably and safely.
- Develop robust Python scripts and tooling to automate routine maintenance, data backups, scaling operations, and system recovery processes.
- Design and enhance the observability stack to provide deep, real-time insights into system health, managing and scaling tools including Prometheus, Grafana, ELK/EFK stack, AWS CloudWatch, and GCP Operations Suite.
- Spearhead reliability initiatives critical to a scaling SaaS platform, driving rigorous capacity planning exercises to stay ahead of growth.
- Own the incident response lifecycle, facilitate blameless postmortems to extract actionable learnings, and define, track, and enforce SLIs, SLOs, and SLAs to ensure the platform consistently meets its reliability guarantees.
- Act as an embedded reliability advocate, collaborating closely with software engineers early in the development lifecycle to ensure applications are designed for deployability, scalability, and resilience.
- Proactively identify system bottlenecks and architectural weaknesses, contribute to process improvements, build internal developer tooling, and maintain comprehensive documentation to elevate team productivity and system understanding.
Requirements
- 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, preferably within a fast-paced SaaS environment.
- Deep, proven proficiency in AWS (EC2, EKS, RDS, VPC, IAM, S3) AND GCP (GKE, Compute Engine, Cloud SQL, IAM, Cloud Storage), with the ability to navigate and optimize multi-cloud architectures.
- Expert-level knowledge of Docker and Kubernetes, including advanced deployment strategies and lifecycle management.
- Strong programming skills in Python.
- Extensive experience with Terraform.
- Hands-on expertise building dashboards and alerting systems using Prometheus, Grafana, and log aggregation stacks (ELK/EFK).
- Solid understanding of cloud networking (VPC peering, load balancing, DNS) and zero-trust security principles in a containerized environment.
Benefits
- Remote-friendly: Drivetrain brings together the best and the brightest, no matter where they are, and provides a great degree of autonomy. We trust our people.
- Open & transparent: Access to all the information needed for creators to do their best work.
- Idea-friendly: Environment to explore new ideas, take risks, make mistakes, and learn, with the best ideas winning regardless of origin.
- Customer-centric: Product-led growth strategy with continuous learning from customers and collaboration to build amazing software.
- Great culture for employees to thrive in and be happy.
Additional details
- Drivetrain is a remote-first company headquartered in the San Francisco Bay Area.
- Founded in 2021 by a couple of ex-Googlers, Drivetrain is a fast-growing company on a trajectory for success with backing from leading venture capital firms.
- Originally posted on Himalayas.