Performance Engineer, Kernels
Sarvam AI
Bengaluru, IndiahybridPosted 28 days ago
S
Skill Required
Infrastructure EngineeringwrittenGitNode.jsGenerative AIMachine LearningFulltime
Key highlights
- Required experience: 5+ years in ML systems with 2+ years authoring production CUDA kernels.
- Notable requirement: Must have a kernel in production that beat the prior baseline by a measurable margin.
- Notable requirement: PTX at a debug-and-modify level — must have inserted hand-written PTX where the compiler missed.
- Must own the kernel layer and author custom CUDA, DSL-based, and PTX kernels where stock libraries leave performance on the table.
- Open-source kernel contributions (e.g., FlashAttention, CUTLASS examples, vLLM/SGLang kernels, DeepEP, Mooncake, or non-trivial Triton) are cited as the highest-yield signal for this role.
- The team works across H100 / H200 / B200 on a multi-node, multi-tenant fleet serving small and large LLMs, MoE, streaming Indic ASR, and multimodal models.
Role overview
This is a Performance Engineer, Kernels role on Sarvam's Performance Engineering team. The kernels team authors the µs-level GPU code, while the inference team integrates it into a running serving stack and owns the system-level numbers. You will own the kernel layer, authoring custom CUDA, DSL-based, and PTX kernels where stock libraries like cuBLAS, cuDNN, FlashAttention, and out-of-the-box Triton leave performance on the table.
Responsibilities
- Part of Sarvam's Performance Engineering team, working alongside a companion Inference role; the two form a vertical stack where the kernels team authors µs-level GPU code and the inference team integrates it into a running serving stack and owns the system-level numbers.
- Own the kernel layer for the team.
- Author custom CUDA, DSL-based, and PTX kernels where stock libraries (cuBLAS, cuDNN, FlashAttention, out-of-the-box Triton) leave performance on the table, closing the gap.
- Ensure that when code lands, production p99 moves, and own the explanation of why.
- Work at the intersection of the serving runtime, the kernel layer, and the SRE org that keeps the fleet alive.
- Own the numbers the rest of the company plans against: how fast we serve, how much it costs, and how much we get out of every GPU.
- Serve multiple model families — small and large LLMs, Mixture-of-Experts, streaming Indic ASR, and multimodal — across a multi-node, multi-tenant fleet on H100 / H200 / B200.
Requirements
- Level: Senior.
- 5+ years in ML systems, with 2+ years authoring production CUDA kernels.
- Have a kernel in production that beat the prior baseline by a measurable margin.
- CUDA at kernel-authoring level: thread-block sizing, shared-memory layout, warp primitives, async copies (cp.async, TMA), and MMA selection.
- CUTLASS / CuTe DSL at a modify-and-extend level, with comfort in the layout algebra.
- PTX at a debug-and-modify level — have inserted hand-written PTX where the compiler missed.
- Nsight Compute and Systems fluency: able to read a roofline plot and propose the fix.
- Attention kernels: have authored or modified at least one (FlashAttention-family, paged, MLA, sliding-window, or sparse).
- Multi-architecture awareness: know what changes from Hopper to Blackwell (TMA, WGMMA, tcgen05).
Nice to have
- Communication kernels — NCCL / NVSHMEM authoring, custom collectives, expert-parallel dispatch (DeepEP-style), AFD bipartite comms (StepMesh-style), KV transfer (DualPath / Mooncake). Strongly desired; dedicated comms-specialist headcount is expected later.
- Open-source kernel contributions — FlashAttention, CUTLASS examples, vLLM / SGLang kernels, DeepEP, Mooncake, or non-trivial Triton work. For this role, the GitHub filter is the highest-yield signal.
- tcgen05, TMA, CTA-cluster launch and distributed shared memory, async pipelining, and the FP4/microscaling paths.
- Grace-side host-path optimization on GH200 / GB200.
Additional details
- The team is hiring two specialized performance roles — Kernels (this posting) and Inference (companion posting).
- If your depth genuinely spans both the kernels and inference roles, apply to either and tell them; most candidates are strongest in one, and they hire for that depth.
- Location: [Bengaluru / Chennai / Hybrid / On-site].
- Team: Performance Engineering.
- Hiring two specialized performance roles (this Kernels posting and a companion Inference posting).
- This is described as a hard, narrow, high-leverage role; they hire engineers who have shipped kernels that beat published baselines on real workloads, not engineers who have used kernels.