Embedded Data Scientist, Chanakya
Sarvam AI
Delhi, IndiaPosted 4 months ago
S
Skill Required
EngineeringData ScientistEmbedded CMachine LearningData StructuresData WarehousingSystem DesigndesigningbuildingPythonPandasdesignNumPyDesign PatternsNLPGenerative AIAIFulltime
Key highlights
- 2–5 years of experience required
- High ownership and impact from day one
- Work on population-scale AI problems for India
- Backed by Lightspeed, Peak XV, and Khosla Ventures
Role overview
Embedded Data Scientists at Sarvam transform complex client data into structures that AI systems can reliably reason over. Deployed alongside Strategic Deployment Engineers at client sites, they work directly with heterogeneous, multimodal data (documents, images, audio, geospatial data, and structured records) to design semantic structures, ontologies, tagging systems, and knowledge graphs that enable effective AI reasoning. The role involves defining data representation, ensuring data quality, and collaborating with cross-functional teams to operationalize data pipelines and improve system performance in real-world deployment conditions.
Responsibilities
- Understand the client's data landscape across documents, imagery, audio, geospatial data, and structured records — including data sources, formats, workflows, and domain terminology
- Design domain ontologies representing entities, relationships, hierarchies, and operational concepts within the client's data environment
- Define document segmentation and chunking strategies that preserve semantic meaning and support effective retrieval
- Work with heterogeneous datasets and define how different modalities should be indexed, embedded, and linked
- Collaborate with Strategic Deployment Engineers to translate semantic structures into operational data ingestion pipelines
- Evaluate how well the AI system retrieves and reasons over client data, and refine structures to improve performance
- Collaborate with the models and other teams to define benchmarks and evaluation criteria that reflect real-world deployment conditions
- Translate insights from client data environments into structured signals for product and engineering teams
- Own the quality of the data layer in assigned accounts, ensuring the system is built on a foundation that enables reliable reasoning at scale
- Work with classified or operationally sensitive datasets in environments where standard tooling may not exist
Requirements
- 2–5 years in data science, applied machine learning, or large-scale data analysis roles
- Strong Python skills including pandas, NumPy, and modern NLP or LLM tooling
- Solid grounding in ML fundamentals — enough to understand model behaviour, contribute to evaluation design, and collaborate with a models team on training and benchmarking
- Experience working with large unstructured datasets including documents, transcripts, reports, or operational records
- Familiarity with LLM-based systems, retrieval pipelines, or vector search systems
- Experience designing or working with data schemas, metadata frameworks, entity models, or semantic data structures
- Comfortable operating with autonomy in client environments; able to do rigorous work without a data team around
- Able to move fluently between domain understanding, data modelling, and AI system design
- Able to move between the technical and the operational: understanding what the data means in the context of what operators actually do with it
Nice to have
- Familiarity with knowledge graphs, ontologies, or semantic data modelling
- Experience with multimodal datasets (text, imagery, audio, geospatial, or structured data)
- Experience operating in constrained or air-gapped environments
- You've worked with real-world, messy, unstructured data and built something rigorous from it
- You are comfortable designing structure where none exists — defining schemas, ontologies, and metadata frameworks from scratch
- You can translate complex data insights into explanations that engineers and client stakeholders can act on
Benefits
- Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar
- High ownership and high impact, from day one
- Everything we do is AI-first, from the way we build and ship to the way we think about problems
- Work on problems that could change how an entire country learns, works, and communicates
Additional details
- Sarvam is building the bedrock of Sovereign AI for India. The company is developing India's full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India.
- Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures.
- Sarvam partners with India's leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.
- Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers of AI with real population-scale impact.