Machine Learning Engineer, Vision
Sarvam AI
Bengaluru, IndiaPosted 4 months ago
S
Skill Required
EngineeringMachine Learning EngineerMachine LearningOWASPStatisticsdebuggingbuildingPyTorchPythondesignGitDesign PatternsAIFulltime
Key highlights
- Undergraduate degree in a technical discipline (CS, statistics, physics, or equivalent) required
- Strong Python and PyTorch expertise required
- Hands‑on experience training or fine‑tuning large vision‑language models
- Work on the full lifecycle of vision‑language model development, from data pipelines to production
- High ownership and high impact from day one
- Collaborate with leading Indian brands such as Tata Capital, SBI Life, CRED, IDFC, and LIC
Role overview
Sarvam is building India's full‑stack sovereign AI platform, developing across research, models, infrastructure, and applications with a focus on AI that works for India. The role involves working across the full lifecycle of vision‑language model development—data pipelines, training, evaluation, and production—while collaborating with clients to deliver end‑to‑end solutions.
Responsibilities
- Design and run training and fine-tuning pipelines for large vision-language models on GPU clusters
- Build multimodal data pipelines — ingestion, filtering, deduplication, synthetic generation, and quality assurance
- Implement and experiment with new architectures and training techniques from research
- Build evaluation harnesses, benchmarks, and automated regression tracking
- Optimise models for inference — quantisation, batching, and serving infrastructure
- Build robust pipelines and integrations that put vision model capabilities in the hands of end users
- Translate real-world problems into well-scoped ML tasks with the right data and evaluation strategy
- Work directly with clients to understand their use cases — document processing, visual search, form extraction — and own the solution end to end
- Build production-grade systems on top of Sarvam Vision and open-source models: multimodal pipelines, retrieval-augmented workflows, and structured output extraction
- Debug and improve deployed solutions — latency, accuracy, edge cases, and integration with client infrastructure
Requirements
- Strong Python and PyTorch — comfortable reading and modifying model internals
- Hands-on experience training or fine-tuning large models, including debugging broken runs
- Experience building data pipelines at scale
- Solid grounding in transformer architectures and modern training techniques
- Comfort with ambiguity — the roadmap is not fully pre-specified
- Strong focus on secure coding practices, code quality, and system reliability
- Undergraduate degree in a technical discipline (CS, statistics, physics, or equivalent)
Nice to have
- Experience with vision-language models or multimodal systems
- Distributed training (FSDP, DeepSpeed, Megatron-LM)
- Post-training methods — RLHF, DPO, or alignment techniques
- Inference optimisation — quantisation, distillation, serving
- Prior exposure to vision-based AI systems or document processing pipelines
- Contributions to open-source projects or a solid GitHub portfolio
Benefits
- Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar
- High ownership and high impact, from day one
- Everything we do is AI-first, from the way we build and ship to the way we think about problems
- You can work on problems that could change how an entire country learns, works, and communicates
- Fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers of AI with real population-scale impact
Additional details
- Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.
- You will work across the full lifecycle of vision-language model (VLM) development — data, training, evaluation, and production. The team's scope will evolve as the field does; we want engineers who are comfortable with that.