Data Engineer
WebEngage
Skill Required
Key highlights
- Experience: 1–3 years professional experience
- Key Benefit: Macbook for Engagers
- Requirement: Strong SQL and Python scripting
- Requirement: Bachelor's degree in CS, Engineering, Math, Stats, or related
Role overview
WebEngage is an enterprise-grade customer engagement and retention platform that helps global brands turn data into measurable revenue impact. Their full-stack retention operating system combines a Customer Data Platform (CDP), real-time behavioral segmentation, omnichannel journey orchestration, AI-driven personalization, and deep analytics. As a Data Engineer, you will own the end-to-end lifecycle of data—from ingesting raw event streams and third-party feeds to building well-modelled, highly reliable datasets that power dashboards, predictive models, and customer journey orchestration. This high-impact role works at the intersection of product, analytics, and backend engineering, requiring a focus on data quality, pipeline observability, cost efficiency, and long-term maintainability.
Responsibilities
- Design, build, and maintain production-grade ETL/ELT pipelines that ingest data from APIs, databases, event streams (Kafka/Pub-Sub), and flat files into the central data warehouse.
- Implement idempotent, incremental load patterns with built-in retry logic, dead-letter queues, and SLA-based alerting to ensure zero-data-loss pipelines.
- Own pipeline observability — set up data freshness checks, row-count validations, schema drift detection, and anomaly alerts using tools like Great Expectations or dbt tests.
- Translate business requirements into clean dimensional models (star/snowflake schemas) and maintain a well-documented data catalogue.
- Design slowly changing dimensions (SCD Type 1/2), bridge tables, and fact tables optimised for analytical query patterns.
- Enforce partitioning, clustering, and materialised view strategies to keep warehouse costs under control while maintaining sub-second query performance.
- Write clean, modular, well-tested Python and SQL code. Follow DRY principles, use version control (Git), and participate in peer code reviews.
- Build reusable transformation frameworks using dbt or equivalent tooling, with proper documentation and testing at every layer (staging → intermediate → mart).
- Containerise data services with Docker and automate deployments via CI/CD pipelines (GitHub Actions / GitLab CI).
- Build interactive dashboards and analytical tools using Streamlit, enabling stakeholders to explore metrics, run ad-hoc analyses, and make data-driven decisions without engineering dependency.
- Design and maintain BI layers — semantic models, KPI definitions, and pre-aggregated mart tables that serve as the single source of truth for reporting across teams.
- Translate raw data into compelling visual narratives using libraries like Plotly, Matplotlib, or Altair; present findings to both technical and non-technical audiences.
- Partner with product managers, analysts, and data scientists to understand data needs and proactively identify gaps in current data coverage.
- Document data lineage, transformation logic, SLAs, and known limitations in a shared knowledge base to enable self-service analytics.
- Contribute to internal engineering guilds, knowledge-sharing sessions, and post-incident reviews for pipeline failures.
Requirements
- Strong SQL skills with expertise in complex queries, performance optimization, and cost-efficient design on cloud data warehouses (BigQuery, Redshift).
- Strong Python scripting for data ingestion, transformation, and validation, with hands-on experience in Pandas, SQLAlchemy, APIs, and automation.
- ETL/ELT- End-to-end ownership of data pipelines, including ingestion, transformation, and loading, with understanding of incremental loads, and backfills.
- Data Modelling- Ability to design dimensional and transactional data models (star/snowflake, SCDs) and translate business needs into optimized table structures.
- Bachelor’s degree in Computer Science, Engineering, Mathematics, Statistics, or a related quantitative field (or equivalent practical experience).
- 1–3 years of professional experience in data engineering, analytics engineering, or a backend role with significant data pipeline work.
- Strong understanding of data warehouse architecture — know when to use wide denormalised tables vs. normalised models, and the trade-offs of each.
- Familiarity with version control workflows (Git branching strategies, pull requests, code reviews) and agile development practices.
- A data quality mindset — you instinctively validate assumptions, add assertions to pipelines, and treat silent data failures as critical incidents.
Nice to have
- Airflow, dbt, Docker, CI/CD
- GCP/AWS, data warehousing concepts
- BI tools / Streamlit / visualization
Benefits
- Full view of team goals and top-level view via monthly & quarterly town hall meetings.
- Highly inclusive work culture that promotes a relaxed, creative and productive environment.
- Autonomy, open communication, and growth opportunities with work-life balance.
- Cutting-edge tools and mentorship (Macbook for Engagers!).
- Best in class medical insurance (with Covid Care facilities).
- Programs for taking care of your mental health.
- Contemporary Leave Policy (beyond sick leaves).
Additional details
- WebEngage is trusted by 800+ brands globally with a presence in India, UAE, KSA, SEA, Europe and beyond.
- WebEngage BLACK is their AI-native layer that brings Agentic capabilities to engagement.
- WebEngage aims to be an equal opportunity employer regardless of race, color, religion, sex, sexual orientation, gender identity, national origin, disability, Veteran status, or any other protected characteristics.
- Originally posted on Himalayas