Skip to main content
Zobhira
Home
Jobs
Certifications
Zobhira
JobsCertificationsTodayAbout
Log inSign up
Zobhira

New job and contest openings, updated every morning on one searchable board.

Find work

  • All jobs
  • Fresher roles
  • Remote roles
  • Certifications

Compete

  • Added today

Popular cities

  • India
  • Bangalore, Karnataka, India
  • Hyderabad, Telangana, India
  • Pune, Maharashtra, India
  • Chennai, Tamil Nadu, India
  • Mumbai, Maharashtra, India

Company

  • About
  • Contact
  • Privacy
  • Terms

Stay updated

One email a week with new roles.

Secure infrastructure
Free to use, no account needed to search

Board updated daily · © 2026 Zobhira. All rights reserved.

Privacy PolicyTerms of Service
Home / Jobs / Sarvam AI

Platform Engineer - AI Infrastructure

Sarvam AI

Bengaluru, IndiaPosted 2 months ago
S

Skill Required

Infrastructure AI EngineerPlatform EngineerInfrastructure EngineerAINginxObservabilityEngineeringKubernetesnetworkingTerraformFirmwarebuildingwrittenPythondesignCloudAPIsSassNode.jsGoMachine LearningFulltime

Key highlights

  • 5+ years of infrastructure/platform software experience required
  • Kubernetes controller/internals-level expertise required
  • High-impact role with end-to-end ownership
  • Work on population-scale AI problems for India

Role overview

Sarvam is building India’s full-stack sovereign AI platform, focusing on research, models, infrastructure, and applications to make AI work for India. The Platform Engineer - AI Infrastructure role involves building the platform that manages a large, multi-vendor GPU fleet, enabling ML teams to use thousands of GPUs without manual intervention for every job. This is a software engineering-heavy role, designing and shipping control-plane services, scheduler integrations, autoscaling controllers, inference-serving platforms, RBAC, quota systems, observability, cost tooling, and APIs/CLIs for ML engineers. The role emphasizes treating the platform as a product and its internal users as customers, with a focus on systems engineering, Kubernetes expertise, and GPU-specific constraints like MIG, gang scheduling, and topology-aware placement.

Responsibilities

  • Build the serving platform: control plane that turns model artifacts into scalable, multi-tenant endpoints, including intelligent routing, load balancing, rollout machinery (canary, blue-green, rollback), traffic splitting, and model-version integration.
  • Develop the scaling and elasticity layer: autoscaling for training (elastic/gang scaling, scale-to-fit) and serving (queue-depth and utilization-driven), capacity pooling across clusters, burst handling, preemption and reclaim, and efficient bin-packing across GPUs.
  • Design and implement scheduling and orchestration: scheduler layer (Kueue, Volcano, Slurm-on-Kubernetes, or custom controllers) with gang scheduling, priority/preemption, queue fairness, quota enforcement, and topology-aware placement to keep job ranks on the same fabric island.
  • Create multi-tenancy, RBAC, and isolation systems: tenant model and namespacing, RBAC/policy, quota and fair-share enforcement, MIG/MPS/time-slicing as self-service tiers, secrets/credential management, and audit logging.
  • Develop networking components: platform-level network plumbing (CNI configuration/custom components), ingress/service routing, RDMA/SR-IOV exposure into pods, tenant network policy/segmentation, and multi-cluster connectivity.
  • Build observability and cost tooling: metrics, logging, and tracing pipelines as platform features with cardinality managed at fleet scale.
  • Design storage and data path abstractions: parallel filesystem abstractions, caching tiers, volume provisioning, and data-locality-aware placement.
  • Enhance developer experience and self-service: CLI, SDK, APIs, job-submission abstractions, and inner-loop tooling for ML engineers.
  • Implement provisioning and infrastructure-as-code: control plane for reproducible cluster setup (Terraform, Crossplane, operators), node lifecycle/image management, driver/firmware rollout automation, and multi-vendor cluster bring-up.

Requirements

  • 5+ years of experience building infrastructure or platform software, with a track record of designing and shipping systems (services/control planes) that others build on.
  • Strong software engineering skills in Go or Python, with the ability to build maintainable systems and debug them in production.
  • Deep Kubernetes expertise at the controller and internals level, including writing operators/controllers, understanding scheduler/API machinery, and knowing where its abstractions leak.
  • Working literacy in GPU-specific platform constraints: MIG and GPU sharing, gang scheduling, topology- and fabric-aware placement, and the contention between training and serving workloads.
  • Product mindset toward internal users: designing APIs/abstractions people want to use, measuring platform success by adoption and self-service rather than tickets closed.
  • Ability to own a capability end to end, from design through rollout to documentation for self-service.

Nice to have

  • Experience building a serving, inference, or training platform (routing, autoscaling, rollout for model endpoints at scale, fine-tuning as a service, etc.).
  • Experience with GPU scheduling systems (Kueue, Volcano, Slurm, Run:ai, or custom schedulers) in multi-tenant production.
  • Multi-tenant isolation (MIG, MPS, time-slicing) shipped as a self-service platform capability.
  • Deep Kubernetes networking experience: CNI internals, custom network components, or RDMA/SR-IOV in pods.
  • On-premise GPU platform work, including multi-vendor or Indian NCP environments.
  • Open-source contributions to Kubernetes, scheduling, or GPU-platform projects.

Benefits

  • Work in a fast-moving, high talent-density team building full-stack AI for India with population-scale impact.
  • Collaborate with researchers, engineers, builders, and business leaders who move fast and hold each other to a high bar.
  • High ownership and high impact from day one.
  • AI-first approach in all aspects of building, shipping, and problem-solving.
  • Opportunity to work on problems that could change how an entire country learns, works, and communicates.

Additional details

  • Sarvam is building the bedrock of Sovereign AI for India, developing India’s full-stack sovereign AI platform.
  • Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures.
  • Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.
  • This is the build side of infrastructure, not the operate side. SREs keep the fleet reliable and carry the pager; this role builds systems that make the fleet usable and reduce breakages.
  • The platform serves two demanding workloads on the same physical infrastructure: training jobs spanning hundreds of GPUs and inference services with flat p99 under production load.

Similar jobs open now

Twilio logo
remote
Digital Marketing Manager
Twilio
Remote - India
View details
ABB logo
Senior Project Engineer -SIS
ABB
Bangalore, Karnataka, India
View details
Cognyte logo
Product Manager
Cognyte
Pune, Maharashtra, India
View details
IBM logo
Data Engineer-Data Platforms-AWS
IBM
Pune, Maharashtra, India
View details
Apply now
LocationBengaluru, India
TypeFulltime
Posted6/29/2026
Apply byOpen

Links are checked every day. If this one stops working, tell us and we'll pull it.

More like this

View all
Twilio logo
Digital Marketing Manager
Twilio
ABB logo
Senior Project Engineer -SIS
ABB
Cognyte logo
Product Manager
Cognyte
Apply now