We are looking for a Data Engineer to design, build, and maintain robust data pipelines and infrastructure that power our analytics, product, and machine learning systems. You'll work across the full data lifecycle: ingestion, transformation, storage, and quality, including datasets used to train and evaluate ML models such as Computer Vision systems and LLM evals. This role suits someone who enjoys building reliable, scalable data systems and cares deeply about data quality and correctness.
Responsibilities
- Design, build, and maintain scalable ETL/ELT pipelines to ingest, clean, transform, and load data from diverse sources (databases, APIs, event streams, files, third-party vendors).
- Build and own data models, schemas, and warehouses/lakehouses that support analytics, reporting, and ML use cases.
- Design and maintain datasets used for ML training and evaluation, including image/video datasets for Computer Vision models and multimodal benchmark sets for LLM evals.
- Implement data quality checks, validation, deduplication, and monitoring to ensure accuracy, completeness, and reliability across pipelines.
- Optimise data storage, partitioning, and query performance for large-scale structured and unstructured data (batch and streaming).
- Build and manage data annotation/labeling workflows for ML datasets, including integration with in-house or third-party labelling tools and vendors.
- Own dataset and pipeline versioning, lineage tracking, and reproducibility practices.
- Collaborate with data scientists, ML engineers, analysts, and product teams to understand data needs and deliver reliable, well-documented datasets.
- Build and maintain orchestration workflows (e. g., Airflow, Dagster, Prefect) and monitor pipeline health, SLAs, and failures.
- Contribute to data governance practices access controls, retention policies, and documentation of schemas and lineage.
Requirements
- 3+ years of experience in data engineering, building and operating production data pipelines at scale.
- Strong Python and SQL skills, with hands-on experience in data pipeline/orchestration frameworks (e. g., Airflow, Dagster, Prefect, or similar).
- Experience with distributed data processing frameworks (Spark, Ray, Dask, or similar) and both batch and streaming data processing.
- Solid understanding of data modelling, warehousing, and lakehouse concepts (e. g., dimensional modelling, partitioning, schema design).
- Experience with cloud data platforms and storage (AWS/GCP/Azure, e. g., S3/GCS, Redshift/BigQuery/Snowflake).
- Experience building or maintaining datasets for ML training/evaluation, including working with unstructured data such as images, video, or text.
- Strong grasp of data quality practices: validation, testing, monitoring, and handling of data drift or anomalies.
- Ability to write clean, well-documented, production-grade code and collaborate effectively with cross-functional teams.
Nice to have
- Experience building datasets specifically for Computer Vision model training (image/video formats, annotation formats like COCO, sharding, streaming loaders).
- Experience building or contributing to LLM evaluation datasets/benchmarks (multimodal or text-only).
- Familiarity with data labeling/annotation tools (e. g., Label Studio, CVAT, Scale AI, Labelbox).
- Experience with data quality/observability tooling and dataset versioning tools (e. g., DVC, LakeFS).
- Experience with infrastructure-as-code and CI/CD for data pipelines.
- Contributions to open-source data or ML tooling.
Additional details
- This job was posted by Harshita Arora from Roadzen.