Cummins India
Formatting this description...
The Data Scientist solves complex analytical and business problems using quantitative approaches through a combination of analytical, mathematical, and technical skills. This role researches, designs, implements, and validates advanced algorithms and machine learning models to analyze diverse data sources and generate actionable insights. The Data Scientist leverages statistical analysis, predictive modeling, machine learning, and Generative AI (GenAI) techniques to drive business outcomes, optimize processes, and enable data-driven decision-making across the organization. Key Responsibilities Leverage data science methodologies to solve complex business challenges and deliver measurable business impact. Research, design, develop, and deploy machine learning, deep learning, predictive analytics, and Generative AI solutions, including Large Language Model (LLM)-based use cases. Build analytical models using statistical methodologies and programming tools to generate insights and recommendations. Perform data extraction, cleaning, preparation, profiling, and transformation to enable high-quality analysis and model development. Conduct feature engineering, model experimentation, hyperparameter tuning, and lifecycle optimization to improve model performance and scalability. Develop descriptive, predictive, and prescriptive models for regression, classification, anomaly detection, forecasting, and pattern recognition. Partner with business stakeholders, product teams, and domain experts to understand requirements and validate model outcomes. Implement model deployment and production pipelines using APIs, containerization, and cloud-based ML platforms. Monitor model performance, accuracy, and drift; recommend improvements and retraining strategies as needed. Communicate methodologies, findings, insights, and recommendations clearly to technical and non-technical stakeholders. Stay current with emerging technologies, trends, and best practices in data science, AI, machine learning, and cloud computing. Contribute to MLOps practices including CI/CD, testing, automation, and governance to ensure reliable production deployments. Skills Strong analytical thinking and problem-solving skills with the ability to work on complex datasets and ambiguous business problems. Effective communication skills with the ability to present technical findings to diverse audiences. Customer-focused mindset with the ability to build strong stakeholder relationships. Strong decision-making capability using data-driven insights. Ability to manage complexity and prioritize high-value opportunities. Strong technical adaptability and passion for emerging technologies. Technical Skills Programming: Advanced proficiency in Python (NumPy, Pandas, scikit-learn) and SQL. Machine Learning & AI: Strong expertise in supervised/unsupervised learning, predictive modeling, deep learning, NLP, and GenAI. Deep Learning Frameworks: Experience with PyTorch, TensorFlow, and/or Keras. LLM & GenAI Tools: Experience with Hugging Face, OpenAI APIs, LangChain, and prompt engineering. Data Mining & Visualization: Ability to extract patterns and insights using statistical analysis and visualization techniques. Statistical Modeling: Strong understanding of probability, hypothesis testing, regression, classification, forecasting, time-series analysis, and anomaly detection. Big Data: Experience with Spark/PySpark and large-scale data processing. Cloud Platforms: Experience with Azure ML, AWS SageMaker, and/or GCP Vertex AI. Model Deployment: Experience deploying models using MLflow, Docker, APIs (FastAPI), or similar frameworks. Software Development: Knowledge of version control, testing frameworks, CI/CD, and secure development practices. Experience Minimum 5+ years of hands-on experience in Data Science, Machine Learning, Artificial Intelligence, or related domains. Proven experience building and deploying end-to-end analytical or machine learning solutions in production environments. Experience processing and managing large, complex datasets. Familiarity with business systems, industry requirements, and data governance regulations. Experience in Agile software development environments. Experience validating and testing machine learning systems for reliability and scalability. Familiarity with CI/CD pipelines and production ML workflows. Nice to Have Experience with MLOps tools and orchestration frameworks such as Airflow or Kubeflow. Experience with modern data platforms such as Databricks or Snowflake. Knowledge of vector databases such as FAISS or Pinecone for GenAI applications. Exposure to RAG (Retrieval-Augmented Generation) architectures and enterprise AI deployment patterns. Experience working with open-source and third-party big data toolsets. Prior internship, co-op, student project, or relevant extracurricular experience in advanced analytics or AI is beneficial for early-career candidates. Competencies Communicates Effectively Customer Focus Decision Quality Manages Complexity Tech Savvy Data Mining Predictive Modeling Programming Requirements Analysis Statistical Modeling Problem Solving Values Differences Qualifications College, university, or equivalent degree in Data Science, Computer Science, Statistics, Mathematics, Engineering, or a related technical discipline required. Advanced degree (Master’s or PhD) in a relevant field is preferred. Equivalent relevant experience may be considered. This position may require licensing for compliance with export controls or sanctions regulations. Job Systems/Information Technology Organization Cummins Inc. Role Category On-site with Flexibility Job Type Exempt - Experienced ReqID 2432032 Relocation Package No 100% On-Site No Show more