Platform Engineer - AI Infrastructure
Sarvam AI
Skill Required
Key highlights
- 5+ years of infrastructure/platform software experience required
- Kubernetes controller/internals-level expertise required
- High-impact role with end-to-end ownership
- Work on population-scale AI problems for India
Role overview
Sarvam is building India’s full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. The Platform Engineer - AI Infrastructure role involves building the platform that manages a large, multi-vendor GPU fleet, enabling ML teams to use thousands of GPUs without manual intervention for every job. This is a software engineering-heavy role, designing and shipping control-plane services, scheduler integrations, autoscaling controllers, inference-serving platforms, RBAC, quota systems, observability, cost tooling, and APIs/CLIs for ML engineers. The role emphasizes treating the platform as a product and its internal users as customers, with a focus on systems engineering, Kubernetes expertise, and GPU-specific constraints like MIG, gang scheduling, and topology-aware placement.
Responsibilities
- Build the serving platform: control plane that turns model artifacts into scalable, multi-tenant endpoints, including intelligent routing, load balancing, rollout machinery (canary, blue-green, rollback), traffic splitting, and model-version integration.
- Develop the scaling and elasticity layer: autoscaling for training (elastic/gang scaling, scale-to-fit) and serving (queue-depth and utilization-driven), capacity pooling across clusters, burst handling, preemption and reclaim, and efficient bin-packing across GPUs.
- Design and implement scheduling and orchestration: scheduler layer (Kueue, Volcano, Slurm-on-Kubernetes, or custom controllers) with gang scheduling, priority/preemption, queue fairness, quota enforcement, and topology-aware placement to keep job ranks on the same fabric island.
- Create multi-tenancy, RBAC, and isolation systems: tenant model and namespacing, RBAC/policy, quota and fair-share enforcement, MIG/MPS/time-slicing as self-service tiers, secrets/credential management, and audit logging.
- Develop networking components: platform-level network plumbing (CNI configuration/custom components), ingress/service routing, RDMA/SR-IOV exposure into pods, tenant network policy/segmentation, and multi-cluster connectivity.
- Build observability and cost tooling: metrics, logging, and tracing pipelines as platform features with cardinality managed at fleet scale.
- Design storage and data path abstractions: parallel filesystem abstractions, caching tiers, volume provisioning, and data-locality-aware placement.
- Enhance developer experience and self-service: CLI, SDK, APIs, job-submission abstractions, and inner-loop tooling for ML engineers.
- Implement provisioning and infrastructure-as-code: control plane for reproducible cluster setup (Terraform, Crossplane, operators), node lifecycle/image management, driver/firmware rollout automation, and multi-vendor cluster bring-up.
Requirements
- 5+ years of experience building infrastructure or platform software, with a track record of designing and shipping systems (services/control planes) that others build on.
- Strong software engineering skills in Go or Python, with the ability to build maintainable systems and debug them in production.
- Deep Kubernetes expertise at the controller and internals level, including writing operators/controllers, understanding scheduler/API machinery, and knowing where its abstractions leak.
- Working literacy in GPU-specific platform constraints: MIG and GPU sharing, gang scheduling, topology- and fabric-aware placement, and the contention between training and serving workloads.
- Product mindset toward internal users: designing APIs/abstractions people want to use, measuring platform success by adoption and self-service rather than tickets closed.
- Ability to own a capability end to end, from design through rollout to documentation for self-service.
Nice to have
- Experience building a serving, inference, or training platform (routing, autoscaling, rollout for model endpoints at scale, fine-tuning as a service, etc.).
- Experience with GPU scheduling systems (Kueue, Volcano, Slurm, Run:ai, or custom schedulers) in multi-tenant production.
- Multi-tenant isolation (MIG, MPS, time-slicing) shipped as a self-service platform capability.
- Deep Kubernetes networking experience: CNI internals, custom network components, or RDMA/SR-IOV in pods.
- On-premise GPU platform work, including multi-vendor or Indian NCP environments.
- Open-source contributions to Kubernetes, scheduling, or GPU-platform projects.
Benefits
- Work in a fast-moving, high talent-density team building full-stack AI for India with population-scale impact.
- Collaborate with researchers, engineers, builders, and business leaders who move fast and hold each other to a high bar.
- High ownership and high impact from day one.
- AI-first approach in all aspects of building, shipping, and problem-solving.
- Opportunity to work on problems that could change how an entire country learns, works, and communicates.
Additional details
- Sarvam is building the bedrock of Sovereign AI for India, developing India’s full-stack sovereign AI platform.
- Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures.
- Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.
- This is the build side of infrastructure, not the operate side. SREs keep the fleet reliable and carry the pager; this role builds systems that make the fleet usable and reduce breakages.
- The platform serves two demanding workloads on the same physical infrastructure: training jobs spanning hundreds of GPUs and inference services with flat p99 under production load.