Back to jobs
Apply on employer siteHybrid

Application Reliability Engineer – Data & AI

MICHELIN

Location

Pune

Salary

Not disclosed

Employment

Other

Experience

5 yrs

Job overview

Complete role details

Location

Pune

Employment type

Other

Workplace

Hybrid

Experience

5 yrs

Seniority

Individual Contributor

Education

Bachelor's degree

Closing date

1 Sept 2026

Requisition ID

R-2025044080

Skills

AzureCI/CDDatabricksGenerative AILLMMachine LearningPythonSpark

Role details

Job description

Application Reliability Engineer – Data & AI

- - - - - - - - - - - -

The Data and AI Application Reliability Engineer you are responsible for owning the post-deployment health, latency, and performance of productionized AI application and Medallion architectures followed for applications.

Acting as the problem-solver for technical issues affecting data pipelines, databases, and deployed AI/ML models, ensuring continuous operation and high user satisfaction.

Partner with core Feature teams to perform root-cause analysis, optimize PySpark jobs, and reduce system debt

Build automated observability dashboards and self-healing mechanisms for automated failure recovery

As a part of Job you are required to balance your responsibilities between Proactive Engineering & Automation and Operational Reliability/ Production Health.

Key Responsibilities

Performance Optimization: Analyze and refactor resource-intensive PySpark jobs, queries, and API endpoints to optimize cost, execution speed, and compute efficiency.

Self-Healing Automation: Develop automated recovery routines, DAG rerun triggers, and data quality checks to minimize manual intervention.

Engineering Alignment: Partner closely with Feature Teams and Architects to establish strict Definition of Done (DoD) standards and production readiness gates for new deployments.

Production Observability & Monitoring: Design, implement, and maintain real-time monitoring and alerting frameworks for Medallion architecture pipelines, feature stores, and AI Applications.

Incident Management & Root-Cause Analysis: Lead technical resolution for high-priority production incidents, conducting thorough post-mortems to eliminate recurring failure patterns.

Technical Expertise:

Azure Data Engineering Stack

Proficiency in Databricks , Python, and PySpark.

Azure (ADLS Gen2, Azure Data Factory, Key Vault, Azure DevOps)

Hands-on experience with Medallion architecture.

Cloud and DevOps Fundamentals

Understanding of cloud computing concepts and Services, specifically Microsoft Azure.

Good handson experience on Python.

Good to Have Technical Abilities :

Understanding of PowerBI reports will be a plus.

DevOps & CI/CD Fundamentals

AI & Machine Learning Fundamentals

Data Science Fundamentals

Generative AI & LLM Fundamentals

Behaviour

Problem Solver: Ability to reverse-engineering complex system behavior and tracking down bugs.

Automation-First: An instinct to automate repetitive tasks .

Clear Communicator: Ability to explain technical root causes to non-technical stakeholders clearly.

Relevant work experience - 5 yrs