Design, optimize, and maintain scalable ETL pipelines using PySpark, and Databricks on cloud platforms (Azure/GCP). Develop automated data validation process to proactively perform data quality checks. Employ key Databricks modules (DeltaLake, Unity Catalog, MLFlow) to facilitate creating and scheduling jobs on Databricks. Optimize the allocation of cloud resources to manage and control cloud cost. Employ GitHub code management repositories and ensure that best practices are being implemented. Build and tune ML/AI and optimization models, identify improvement opportunities, and perform experiments to demonstrate incremental value. Have frequent conversations with Business Stakeholders to understand their requirements and concerns. Explain data deficiencies, model performance/root cause analysis, and model explainability. Follow best practices in Data Architecture, Coding, and Project Management operations. Collaborate with cross-functional teams, such as Customers' Stakeholders, Engagement Managers, Data Ops/Job Monitoring, Product Management & Software Engineering. Expand the use of analytics, ML/AI, mathematical optimization, Gen-AI/LLMs and Agentic-AI in the context of Retail/CPG business use cases such as anomaly detection, demand forecasting, price elasticity modeling, promotions features & strategy simulation, product cannibalization and halo modeling, markdown optimization, product allocation, reorder/replenishment, size and pack optimization, workforce scheduling & task optimization Minimum Education: Master's degree in engineering, computer science, data science, operations research, statistics, mathematics, quantitative sciences or relevant work experience Minimum Work Experience (years): 4+ years of experience in Data Science/Data Engineering with emphasis on the full lifecycle of Data Science-ML/AI projects. Within that timeframe, experience is expected in: Python/PySpark, SQL, and relational or NoSQL databases, and cloud resource management. Key Skills and Competencies: Experience working with AWS, Azure, or GCP cloud environments. Experience implementing advanced analytics, ML/AI algorithms (such as statistical time series: ESMs, ARIMA; machine learning: Random Forests, GBMs; Neural Networks: TiDE & DenseNet, foundational models) and mathematical (constrained linear/non-linear and network) optimization models Proven experience building end-to-end production grade Data & ML/AI pipelines using PySpark, Python (Pandas/NumPy) and SQL. Experience working with Git (or similar code management repositories) as a collaboration tool. Experience with orchestration tools like Databricks, Airflow (or similar tools like Snowflake, Dagster, etc.). Working knowledge of GenAI/LLMs, Agentic-AI and related frameworks (e.g. LangChain) will be a plus Understanding of Retail/CPG industry business challenges with an emphasis on Supply Chain, Pricing and Workforce optimization applications is preferred Excellent verbal and written communication skills, especially as it relates to technical communications. Ability to present technical analysis to business stakeholders. Demonstrated ability to learn new technologies quickly and independently. Ability to work independently with minimal supervision and achieve stretch goals in a very innovative and fast-paced environment. Licenses/Certifications, special qualifications: N/A Equivalencies: Relevant work experience may be substituted for a degree.
💡 Quick Summary
Seeking a career-building opportunity? The Data Scientist (IN), Senior position is now open for candidates interested in the IT Engineer & Developer Jobs sector. This role in Mumbai offers a professional environment and growth potential.
Requirement Snapshot: Candidates should possess basic communication skills, a proactive attitude, and the ability to work in a team. Experience in IT Engineer & Developer Jobs is a plus.