Infrastructure SRE - HPC

Sarvam AI

Bengaluru, IndiaPosted 2 months ago
S

Skill Required

Infrastructure Site Reliability EngineerInfrastructure EngineerObservabilityEngineeringKubernetesnetworkingdebuggingFirmwarebuildingPythonSassNode.jsGoAIMachine LearningFulltime

Key highlights

  • 5+ years of infrastructure/SRE experience, including 2+ years with GPU clusters at scale.
  • Proficiency in Python or Go for building internal tooling.
  • High ownership and high impact from day one.
  • Opportunity to work on AI problems that could change an entire country's learning, work, and communication.

Role overview

Sarvam runs a large, multi‑vendor GPU fleet that serves two demanding workloads on the same physical infrastructure: training jobs that span hundreds of GPUs and must run uninterrupted for weeks, and inference services that must hold a flat p99 under production load. Keeping both healthy at once is a hard, specialized reliability problem, and it is the problem this team exists to solve. This role goes beyond basic Kubernetes administration; it requires deep expertise in parallel filesystems under heavy checkpoint load, RDMA fabrics that degrade quietly, NCCL hangs whose root cause may be the network or the kernel, driver and firmware drift across heterogeneous hardware, and distributed training failures that masquerade as infrastructure faults.

Responsibilities

  • Operate the GPU fleet end to end across training and serving – provisioning, observability, capacity, and fleet health.
  • Hold a meaningful on‑call rotation, write runbooks that hold up under pressure, and drive postmortems that produce durable fixes.
  • Build the internal tooling the team relies on, rather than operating off‑the‑shelf systems alone.
  • Partner with ML and platform teams to keep large runs alive and serving latency predictable.

Requirements

  • 5+ years in infrastructure or site reliability engineering, including 2+ years operating GPU clusters at scale.
  • Demonstrated on‑call ownership of infrastructure that mattered, with a track record of postmortems that led to real change.
  • Proficiency in Python or Go, used to build and maintain internal tooling.
  • Working fluency across all five areas of focus (distributed high‑performance storage, fabric & RDMA networking, GPU systems reliability, Kubernetes platform reliability, training & inference workload reliability) – enough to recognize, triage, and route a problem outside your specialty.
  • Bring depth in one of the five areas below; expect to be conversational across the rest.
  • For the Storage and Fabric areas of focus, deep domain expertise may be weighed against the GPU‑cluster requirement; exceptional specialists with less direct GPU‑fleet time are encouraged to apply.

Nice to have

  • Experience with Slurm and Kubernetes hybrid environments.
  • On‑premise GPU deployment experience, including coordination with datacenter operations on power, cooling, and InfiniBand cabling.
  • Familiarity with Indian NCPs, DGX SuperPOD, Lambda, CoreWeave, NeevCloud, etc.
  • Production‑grade multi‑tenant GPU isolation (MIG, MPS, time‑slicing).

Benefits

  • Fast‑moving, high talent‑density team building full‑stack AI for India with real population‑scale impact.
  • Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar.
  • High ownership and high impact from day one.
  • Everything is AI‑first, from the way we build and ship to the way we think about problems.
  • Opportunity to work on problems that could change how an entire country learns, works, and communicates.

Additional details

  • About Sarvam: building the bedrock of Sovereign AI for India with a full‑stack sovereign AI platform across research, models, infrastructure, and applications.
  • Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures.
  • Partners include Tata Capital, SBI Life, CRED, IDFC, and LIC.
  • This is not a Kubernetes administration role; Kubernetes fluency is assumed as a baseline.
  • When you apply, indicate the area of focus that best matches your experience; strong generalists are welcome and will be placed where their depth is most useful.
  • The team is hiring specialists rather than identical generalists, expecting genuine depth in one focus area and working fluency across the others.
Apply now