Data Engineer

EXL

Pune City, Maharashtra, IndiaPosted 5 days ago
EXL logo

Role tags

Data Engineer

Tech stack mentioned

PythonHadoopSQLAWSLinux

Role overview

Formatting this description...

Data Engineer 📍 Locations: Pune, Gurugram, Hyderabad, Bangalore 🏢 Work Mode: Hybrid 💼 Experience: 5+ Years Position Summary We are looking for a skilled Data Engineer with 5+ years of experience in building, maintaining, and optimizing scalable data pipelines using PySpark, Python, and SQL. The ideal candidate should possess strong hands-on expertise in Big Data technologies such as Hive and Oozie, along with a solid understanding of ETL processes, data processing frameworks, and cloud environments (preferably AWS). The role involves collaborating with cross-functional teams to support enterprise data platforms, analytics initiatives, and business-critical data processing requirements. Key Responsibilities Design, develop, and maintain scalable data pipelines using PySpark, Python, Hive, and Oozie. Develop efficient SQL queries for data extraction, transformation, validation, and reporting. Integrate data from multiple source systems while ensuring accuracy, consistency, and completeness. Monitor, troubleshoot, and optimize data pipelines to improve performance and reliability. Collaborate with business stakeholders, analysts, and engineering teams to understand data requirements and deliver effective solutions. Implement best practices for data engineering, documentation, coding standards, and performance optimization. Support data quality initiatives and ensure adherence to data governance standards. Participate in code reviews, testing, deployment, and production support activities. Required Skills 5+ years of hands-on experience in PySpark, Python, and SQL. Strong experience working with Big Data technologies, including Hive and Oozie. Good understanding of ETL/ELT concepts, data processing, and data integration techniques. Working knowledge of cloud platforms, preferably AWS. Experience handling large-scale datasets and optimizing data processing workflows. Strong analytical and problem-solving skills. Secondary Skills Basic to intermediate experience with Databricks. Familiarity with Linux/Unix commands and environment management on edge nodes. Exposure to workflow scheduling and orchestration tools. Knowledge of monitoring and debugging data pipelines. Good to Have Exposure to the Financial Services/Banking domain. Understanding of Data Warehousing concepts and best practices. Experience with modern data transformation frameworks such as DBT. Exposure to orchestration platforms such as Dagster. Knowledge of cloud-based data engineering architectures and modern data platforms. Technical Skills PySpark, Python, SQL, Hive, Oozie, Hadoop, AWS, Databricks, Linux/Unix, ETL/ELT, Data Warehousing, DBT, Dagster

Apply now