Data Engineer
Fusemachines
IndiaremotePosted 13 days ago
Skill Required
Data-EngineeringData-Pipeline-DevelopmentAWS-Data-EngineeringBackend-EngineeringAdTech-EngineeringData-EngineerData-Engineer-JobsData-Engineering-JobsData-Engineering-PositionsData EngineerQuery OptimizationObservabilityAirflowSparkPythonReactKafkaJavagRPCETLSQLAWSGenerative AIMachine LearningContract
Key highlights
- Remote, Full-time role in Indian Timezone
- 4–8 years of engineering experience required
- AWS-based data stack (S3, Glue, Athena, Airflow) required
- Full remote flexibility for India/Nepal contractors
- Clear growth path to Staff/Principal Engineer roles
- Focus on backend data pipelines and media measurement
Role overview
Fusemachines is a leading AI strategy, talent, and education services provider with a mission to democratize AI, operating in 4 countries with over 400 employees. The role is for a highly analytical Senior/Mid-Senior Data Engineer (full-time contractor, Indian Timezone) to build and scale data systems for a global media measurement and consumer intelligence company. The focus is on data quality assessments, backend data pipelines, and pipeline enrichment (not front-end development), involving AWS-based architectures, troubleshooting, and powering measurement systems for advertising effectiveness and business decisions at scale.
Responsibilities
- Design, build, and maintain robust, scalable data enrichment pipelines using Apache Spark within an AWS environment.
- Utilize AWS S3 as the core target data storage, leveraging AWS Glue and Athena for data cataloging and querying.
- Configure, manage, and monitor automated data workflows and pipelines using Airflow.
- Perform deep analytical assessments on pipeline data to evaluate data quality, ensure completeness, and maintain high security standards.
- Define and enforce data contracts, schemas, and quality standards, implement validation checks and monitoring processes to ensure data accuracy.
- Analyze core data to provide actionable insights and proactively suggest pipeline enhancements or new data applications based on findings.
- Improve performance, scalability, and cost-efficiency of pipelines, queries, and storage services while resolving bottlenecks.
- Act as a technical detective to troubleshoot data delivery and device tracking issues across hybrid server-to-server systems.
- Implement monitoring, observability, and debugging mechanisms across all layers (data → API → client SDKs).
- Collaborate with global product, engineering, and data science teams to thoroughly understand logic requirements and provide backend engineering support.
- Execute data governance strategies encompassing cataloging, lineage tracking, and schema design working closely with the Data Architect.
Requirements
- 4–8 years of engineering experience, with a strong background in real-world data engineering development on AWS.
- Bachelor’s degree in Computer Science or related field.
- Proven experience working with an AWS-based data stack, specifically using S3 as the target data storage, Glue for processing, Athena for querying (MPP), and Airflow for workflow orchestration.
- Strong hands-on experience building data enrichment pipelines using Apache Spark / PySpark.
- Strong analytical background with a verified ability to evaluate data quality, analyze datasets to uncover trends, and suggest potential development applications based on findings.
- Proficiency in Python or Java, combined with a strong understanding of SQL, writing advanced queries, and query optimization.
- Skilled in data integration (ETL/ELT) from diverse sources like APIs (REST/gRPC), flat files, databases, and event streaming.
- Ability to dig into multi-layered architectures to troubleshoot data mismatches or technical tracking issues.
- Great problem-solving skills, high attention to detail, and the ability to define and document data engineering processes and data flows.
Nice to have
- Advanced studies or exposure to AI/ML.
- Familiarity with the digital advertising ecosystem (AdTech), audience segmentation, targeting, and media measurement concepts (reach, frequency, attribution, and campaign performance).
- Exposure to event data pipelines (impressions, clicks, conversions) and privacy-aware/identity-driven data systems.
- Experience with streaming systems (Kafka, etc.).
- Understanding of frontend technologies (React or similar).
- Understanding of mobile SDK development.
- Exposure to AI/ML pipelines or LLM ecosystems.
- Understanding of developer platforms (SDKs, APIs, integration tooling).
Benefits
- Work on global media and consumer measurement pipelines utilized by leading brands worldwide.
- Engage in deep, logic-driven engineering problems where data analysis and quality are prioritized over UI delivery.
- Gain hands-on ownership of an extensive AWS big data infrastructure optimized for high throughput.
- Direct exposure to uniquely rich datasets spanning retail, consumer behavior, and advertising ecosystems.
- Opportunity to solve complex, high-impact problems at the intersection of data, measurement, and AI.
- Collaborate with cross-functional, highly data-driven global teams with a strong focus on innovation and modern engineering practices.
- Clear growth path toward Staff/Principal Engineer and Platform Leadership roles.
- Full remote flexibility for contractors based in India or Nepal (with hybrid options available if located near the Pune office).
- Equal Opportunity Employer: Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.
Additional details
- Type: Remote, Full-time
- Founded by Sameer Maskey Ph.D., Adjunct Associate Professor at Columbia University.
- Presence in 4 countries (Nepal, United States, Canada, and Dominican Republic).
- Originally posted on Himalayas