Skip to main content
Zobhira
Home
Jobs
Certifications
Zobhira
JobsCertificationsTodayAbout
Log inSign up
Zobhira

New job and contest openings, updated every morning on one searchable board.

Find work

  • All jobs
  • Fresher roles
  • Remote roles
  • Certifications

Compete

  • Added today

Popular cities

  • India
  • Bangalore, Karnataka, India
  • Hyderabad, Telangana, India
  • Pune, Maharashtra, India
  • Chennai, Tamil Nadu, India
  • Mumbai, Maharashtra, India

Company

  • About
  • Contact
  • Privacy
  • Terms

Stay updated

One email a week with new roles.

Secure infrastructure
Free to use, no account needed to search

Board updated daily · © 2026 Zobhira. All rights reserved.

Privacy PolicyTerms of Service
Home / Jobs / Sarvam AI

Performance Engineer, Inference

Sarvam AI

Bengaluru, IndiahybridPosted 28 days ago
S

Skill Required

Infrastructure EngineeringdesignC++Node.jsGenerative AIMachine LearningFulltime

Key highlights

  • Required experience: 5+ years in ML systems, with 2+ years on inference serving at production scale
  • Key requirement: Source-level fluency (with modifications) in one of SGLang, vLLM, Dynamo, or TensorRT-LLM, plus hands-on production experience serving 100B+ parameter models across multi-node parallelism
  • Key requirement: Build-and-train competency in speculative decoding — trained own draft models/speculators (e.g., EAGLE / DFlash), distilled from a target, tuned acceptance rate against a live serving distribution
  • Notable benefit/responsibility: Owns the company-facing scoreboard — TTFT (p50/p95/p99), TPOT, throughput, GPU utilization, and cost per million tokens — and on-call ownership of an inference SLO
  • Team context: Senior role on the Performance Engineering team at Sarvam, working at the intersection of the serving runtime, kernel layer, and SRE across a multi-node, multi-tenant fleet of Hoppers and Blackwells
  • Bonus: Upstream contributions to SGLang, vLLM, Dynamo, llm-d, TensorRT-LLM, or LMDeploy on non-trivial code paths

Role overview

Performance Engineer, Inference at Sarvam — part of the Performance Engineering team, sitting between the kernels team (which authors µs-level GPU code) and SRE (which keeps the multi-tenant fleet alive). The role owns Sarvam's production serving path for large distributed models end-to-end: integrating kernel and model artifacts into a multi-node, multi-tenant stack, building and training custom speculators, and producing the latency/throughput/cost numbers the company plans against. Level is Senior; location options are Bengaluru / Chennai / Hybrid / On-site; the team is hiring two specialized roles (Inference and Kernels) as a vertical stack, and most candidates are expected to be strongest in one.

Responsibilities

  • Own Sarvam's production serving path for large distributed models end to end
  • Be source-level fluent in at least one of SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM — read and modify it where stock behavior does not fit our workloads
  • Operate and extend a distributed-serving stack at depth: disaggregated prefill-decode across nodes, distributed KV/cache transfer, and the routing and scheduling that span them
  • Integrate artifacts from the model and kernel teams into a running multi-node, multi-tenant stack
  • Build and train your own speculators — draft models, distillation from the target, acceptance-rate tuning against the live serving distribution — rather than only wiring in stock implementations
  • Produce, and defend, the latency and throughput numbers the company plans against
  • Spend significant time in cross-team work with architecture co-design, the kernels team, the model team, and SRE
  • Be accountable to a scoreboard of: TTFT (p50 / p95 / p99), TPOT, throughput, GPU utilization, and cost per million tokens
  • Own on-call responsibility for an inference SLO

Requirements

  • 5+ years in ML systems, with 2+ years on inference serving at production scale
  • A record showing concrete outcomes — tokens per day, throughput wins, p99 reductions — rather than 'deployed a model'
  • Experience serving 100B+ parameter models in production across multi-node tensor, pipeline, or expert parallelism
  • Source-level fluency in one of SGLang, vLLM, Dynamo, or TensorRT-LLM — having modified the scheduler, the KV allocator, or the disaggregation path
  • Reading-level familiarity with the other three of SGLang, vLLM, Dynamo, TensorRT-LLM
  • Distributed serving at operating-and-extending depth: a disaggregated prefill-decode stack, distributed KV/cache transfer (Mooncake or equivalent), and cross-node routing and scheduling — have run one of these in production and modified it where it didn't fit
  • Speculative decoding as a build-and-train competency: have trained your own draft models or speculators (EAGLE / DFlash or otherwise), distilled them from a target model, measured and tuned acceptance rate against a real serving distribution, and composed speculation with the rest of the stack — not only integrated a stock implementation
  • Deep understanding of KV cache internals: block tables, copy-on-write, prefix sharing, and fragmentation
  • Working command of TP / PP / EP, NCCL primitives, and how they interact with the scheduler
  • Experience with multi-tenant serving: model co-location and MIG / MPS isolation
  • Profiling fluency with Nsight Systems, framework tracing, and py-spy / perf
  • C++ and CUDA at a read-and-modify level

Nice to have

  • Upstream contributions to SGLang, vLLM, Dynamo, llm-d, TensorRT-LLM, or LMDeploy on non-trivial code paths
  • Direct production experience with Dynamo or llm-d at scale
  • Having operated a forked runtime in production
  • Published or shipped speculator work — a draft model or speculative-decoding technique you trained and measured
  • MoE serving at scale, long-context (128K+), multi-model serving, or Indic and multilingual workloads

Additional details

  • Role title: Performance Engineer, Inference
  • Part of Sarvam's Performance Engineering team
  • Hiring two specialized performance roles — Inference (this posting) and Kernels (companion posting)
  • The two roles form a vertical stack: the kernels team authors the µs-level GPU code, and the inference team integrates it into a running serving stack and owns the system-level numbers
  • If a candidate's depth genuinely spans both roles, they may apply to either and note that — but most candidates are strongest in one, and the team hires for that depth
  • Locations: Bengaluru / Chennai / Hybrid / On-site
  • Team: Performance Engineering
  • Level: Senior
  • Sarvam serves multiple model families — small and large LLMs, Mixture-of-Experts, Indic ASR & TTS, streaming models, and multimodal models
  • Workload runs across a multi-node, multi-tenant fleet of Hoppers and Blackwells
  • The Performance Engineering team owns the numbers the rest of the company plans against: how fast we serve, how much it costs, and how much we get out of every GPU
  • The team works at the intersection of the serving runtime, the kernel layer, and the SRE org that keeps the fleet alive

Similar jobs open now

Twilio logo
remote
Digital Marketing Manager
Twilio
Remote - India
View details
ABB logo
Senior Project Engineer -SIS
ABB
Bangalore, Karnataka, India
View details
Cognyte logo
Product Manager
Cognyte
Pune, Maharashtra, India
View details
IBM logo
Data Engineer-Data Platforms-AWS
IBM
Pune, Maharashtra, India
View details
Apply now
LocationBengaluru, India
TypeFulltime
Posted8/10/2026
Apply byOpen

Links are checked every day. If this one stops working, tell us and we'll pull it.

More like this

View all
Twilio logo
Digital Marketing Manager
Twilio
ABB logo
Senior Project Engineer -SIS
ABB
Cognyte logo
Product Manager
Cognyte
Apply now