Skip to main content
Zobhira
Home
Jobs
Certifications
Zobhira
JobsCertificationsTodayAbout
Log inSign up
Zobhira

New job and contest openings, updated every morning on one searchable board.

Find work

  • All jobs
  • Fresher roles
  • Remote roles
  • Certifications

Compete

  • Added today

Popular cities

  • India
  • Bangalore, Karnataka, India
  • Hyderabad, Telangana, India
  • Pune, Maharashtra, India
  • Chennai, Tamil Nadu, India
  • Mumbai, Maharashtra, India

Company

  • About
  • Contact
  • Privacy
  • Terms

Stay updated

One email a week with new roles.

Secure infrastructure
Free to use, no account needed to search

Board updated daily · © 2026 Zobhira. All rights reserved.

Privacy PolicyTerms of Service
Home / Jobs / Sarvam AI

Performance Engineer, Kernels

Sarvam AI

Bengaluru, IndiahybridPosted 28 days ago
S

Skill Required

Infrastructure EngineeringwrittenGitNode.jsGenerative AIMachine LearningFulltime

Key highlights

  • Required experience: 5+ years in ML systems with 2+ years authoring production CUDA kernels.
  • Notable requirement: Must have a kernel in production that beat the prior baseline by a measurable margin.
  • Notable requirement: PTX at a debug-and-modify level — must have inserted hand-written PTX where the compiler missed.
  • Must own the kernel layer and author custom CUDA, DSL-based, and PTX kernels where stock libraries leave performance on the table.
  • Open-source kernel contributions (e.g., FlashAttention, CUTLASS examples, vLLM/SGLang kernels, DeepEP, Mooncake, or non-trivial Triton) are cited as the highest-yield signal for this role.
  • The team works across H100 / H200 / B200 on a multi-node, multi-tenant fleet serving small and large LLMs, MoE, streaming Indic ASR, and multimodal models.

Role overview

This is a Performance Engineer, Kernels role on Sarvam's Performance Engineering team. The kernels team authors the µs-level GPU code, while the inference team integrates it into a running serving stack and owns the system-level numbers. You will own the kernel layer, authoring custom CUDA, DSL-based, and PTX kernels where stock libraries like cuBLAS, cuDNN, FlashAttention, and out-of-the-box Triton leave performance on the table.

Responsibilities

  • Part of Sarvam's Performance Engineering team, working alongside a companion Inference role; the two form a vertical stack where the kernels team authors µs-level GPU code and the inference team integrates it into a running serving stack and owns the system-level numbers.
  • Own the kernel layer for the team.
  • Author custom CUDA, DSL-based, and PTX kernels where stock libraries (cuBLAS, cuDNN, FlashAttention, out-of-the-box Triton) leave performance on the table, closing the gap.
  • Ensure that when code lands, production p99 moves, and own the explanation of why.
  • Work at the intersection of the serving runtime, the kernel layer, and the SRE org that keeps the fleet alive.
  • Own the numbers the rest of the company plans against: how fast we serve, how much it costs, and how much we get out of every GPU.
  • Serve multiple model families — small and large LLMs, Mixture-of-Experts, streaming Indic ASR, and multimodal — across a multi-node, multi-tenant fleet on H100 / H200 / B200.

Requirements

  • Level: Senior.
  • 5+ years in ML systems, with 2+ years authoring production CUDA kernels.
  • Have a kernel in production that beat the prior baseline by a measurable margin.
  • CUDA at kernel-authoring level: thread-block sizing, shared-memory layout, warp primitives, async copies (cp.async, TMA), and MMA selection.
  • CUTLASS / CuTe DSL at a modify-and-extend level, with comfort in the layout algebra.
  • PTX at a debug-and-modify level — have inserted hand-written PTX where the compiler missed.
  • Nsight Compute and Systems fluency: able to read a roofline plot and propose the fix.
  • Attention kernels: have authored or modified at least one (FlashAttention-family, paged, MLA, sliding-window, or sparse).
  • Multi-architecture awareness: know what changes from Hopper to Blackwell (TMA, WGMMA, tcgen05).

Nice to have

  • Communication kernels — NCCL / NVSHMEM authoring, custom collectives, expert-parallel dispatch (DeepEP-style), AFD bipartite comms (StepMesh-style), KV transfer (DualPath / Mooncake). Strongly desired; dedicated comms-specialist headcount is expected later.
  • Open-source kernel contributions — FlashAttention, CUTLASS examples, vLLM / SGLang kernels, DeepEP, Mooncake, or non-trivial Triton work. For this role, the GitHub filter is the highest-yield signal.
  • tcgen05, TMA, CTA-cluster launch and distributed shared memory, async pipelining, and the FP4/microscaling paths.
  • Grace-side host-path optimization on GH200 / GB200.

Additional details

  • The team is hiring two specialized performance roles — Kernels (this posting) and Inference (companion posting).
  • If your depth genuinely spans both the kernels and inference roles, apply to either and tell them; most candidates are strongest in one, and they hire for that depth.
  • Location: [Bengaluru / Chennai / Hybrid / On-site].
  • Team: Performance Engineering.
  • Hiring two specialized performance roles (this Kernels posting and a companion Inference posting).
  • This is described as a hard, narrow, high-leverage role; they hire engineers who have shipped kernels that beat published baselines on real workloads, not engineers who have used kernels.

Similar jobs open now

ABB logo
Senior Project Engineer -SIS
ABB
Bangalore, Karnataka, India
View details
Cognyte logo
Product Manager
Cognyte
Pune, Maharashtra, India
View details
IBM logo
Data Engineer-Data Platforms-AWS
IBM
Pune, Maharashtra, India
View details
IBM logo
Technical Consultant-AI Integration
IBM
Pune, Maharashtra, India
View details
Apply now
LocationBengaluru, India
TypeFulltime
Posted8/10/2026
Apply byOpen

Links are checked every day. If this one stops working, tell us and we'll pull it.

More like this

View all
ABB logo
Senior Project Engineer -SIS
ABB
Cognyte logo
Product Manager
Cognyte
IBM logo
Data Engineer-Data Platforms-AWS
IBM
Apply now