Performance Engineer, Inference
Sarvam AI
Bengaluru, IndiahybridPosted 28 days ago
S
Skill Required
Infrastructure EngineeringdesignC++Node.jsGenerative AIMachine LearningFulltime
Key highlights
- Required experience: 5+ years in ML systems, with 2+ years on inference serving at production scale
- Key requirement: Source-level fluency (with modifications) in one of SGLang, vLLM, Dynamo, or TensorRT-LLM, plus hands-on production experience serving 100B+ parameter models across multi-node parallelism
- Key requirement: Build-and-train competency in speculative decoding — trained own draft models/speculators (e.g., EAGLE / DFlash), distilled from a target, tuned acceptance rate against a live serving distribution
- Notable benefit/responsibility: Owns the company-facing scoreboard — TTFT (p50/p95/p99), TPOT, throughput, GPU utilization, and cost per million tokens — and on-call ownership of an inference SLO
- Team context: Senior role on the Performance Engineering team at Sarvam, working at the intersection of the serving runtime, kernel layer, and SRE across a multi-node, multi-tenant fleet of Hoppers and Blackwells
- Bonus: Upstream contributions to SGLang, vLLM, Dynamo, llm-d, TensorRT-LLM, or LMDeploy on non-trivial code paths
Role overview
Performance Engineer, Inference at Sarvam — part of the Performance Engineering team, sitting between the kernels team (which authors µs-level GPU code) and SRE (which keeps the multi-tenant fleet alive). The role owns Sarvam's production serving path for large distributed models end-to-end: integrating kernel and model artifacts into a multi-node, multi-tenant stack, building and training custom speculators, and producing the latency/throughput/cost numbers the company plans against. Level is Senior; location options are Bengaluru / Chennai / Hybrid / On-site; the team is hiring two specialized roles (Inference and Kernels) as a vertical stack, and most candidates are expected to be strongest in one.
Responsibilities
- Own Sarvam's production serving path for large distributed models end to end
- Be source-level fluent in at least one of SGLang, vLLM, NVIDIA Dynamo, or TensorRT-LLM — read and modify it where stock behavior does not fit our workloads
- Operate and extend a distributed-serving stack at depth: disaggregated prefill-decode across nodes, distributed KV/cache transfer, and the routing and scheduling that span them
- Integrate artifacts from the model and kernel teams into a running multi-node, multi-tenant stack
- Build and train your own speculators — draft models, distillation from the target, acceptance-rate tuning against the live serving distribution — rather than only wiring in stock implementations
- Produce, and defend, the latency and throughput numbers the company plans against
- Spend significant time in cross-team work with architecture co-design, the kernels team, the model team, and SRE
- Be accountable to a scoreboard of: TTFT (p50 / p95 / p99), TPOT, throughput, GPU utilization, and cost per million tokens
- Own on-call responsibility for an inference SLO
Requirements
- 5+ years in ML systems, with 2+ years on inference serving at production scale
- A record showing concrete outcomes — tokens per day, throughput wins, p99 reductions — rather than 'deployed a model'
- Experience serving 100B+ parameter models in production across multi-node tensor, pipeline, or expert parallelism
- Source-level fluency in one of SGLang, vLLM, Dynamo, or TensorRT-LLM — having modified the scheduler, the KV allocator, or the disaggregation path
- Reading-level familiarity with the other three of SGLang, vLLM, Dynamo, TensorRT-LLM
- Distributed serving at operating-and-extending depth: a disaggregated prefill-decode stack, distributed KV/cache transfer (Mooncake or equivalent), and cross-node routing and scheduling — have run one of these in production and modified it where it didn't fit
- Speculative decoding as a build-and-train competency: have trained your own draft models or speculators (EAGLE / DFlash or otherwise), distilled them from a target model, measured and tuned acceptance rate against a real serving distribution, and composed speculation with the rest of the stack — not only integrated a stock implementation
- Deep understanding of KV cache internals: block tables, copy-on-write, prefix sharing, and fragmentation
- Working command of TP / PP / EP, NCCL primitives, and how they interact with the scheduler
- Experience with multi-tenant serving: model co-location and MIG / MPS isolation
- Profiling fluency with Nsight Systems, framework tracing, and py-spy / perf
- C++ and CUDA at a read-and-modify level
Nice to have
- Upstream contributions to SGLang, vLLM, Dynamo, llm-d, TensorRT-LLM, or LMDeploy on non-trivial code paths
- Direct production experience with Dynamo or llm-d at scale
- Having operated a forked runtime in production
- Published or shipped speculator work — a draft model or speculative-decoding technique you trained and measured
- MoE serving at scale, long-context (128K+), multi-model serving, or Indic and multilingual workloads
Additional details
- Role title: Performance Engineer, Inference
- Part of Sarvam's Performance Engineering team
- Hiring two specialized performance roles — Inference (this posting) and Kernels (companion posting)
- The two roles form a vertical stack: the kernels team authors the µs-level GPU code, and the inference team integrates it into a running serving stack and owns the system-level numbers
- If a candidate's depth genuinely spans both roles, they may apply to either and note that — but most candidates are strongest in one, and the team hires for that depth
- Locations: Bengaluru / Chennai / Hybrid / On-site
- Team: Performance Engineering
- Level: Senior
- Sarvam serves multiple model families — small and large LLMs, Mixture-of-Experts, Indic ASR & TTS, streaming models, and multimodal models
- Workload runs across a multi-node, multi-tenant fleet of Hoppers and Blackwells
- The Performance Engineering team owns the numbers the rest of the company plans against: how fast we serve, how much it costs, and how much we get out of every GPU
- The team works at the intersection of the serving runtime, the kernel layer, and the SRE org that keeps the fleet alive