Regeneron logo

AI/MLOps SRE Lead Engineer

Regeneron

Hybrid
Hyderabad
Full-time
Lead
Senior
6+ yrs
Salary not listedPosted 1h ago

Real job — pulled straight from Regeneron’s careers page · Verified September 28, 2026 · No reposts.

Job description

Regeneron is hiring a AI/MLOps SRE Lead Engineer — a full-time, based in Hyderabad role. Apply directly on Regeneron's careers page below.

AI/MLOps SRE Lead Engineer

Location: Hyderabad

Remote Type: Hybrid

Time Type: Full time

Job Description

Build our future together

Regeneron is founded on the belief that the right idea, combined with the right team, can lead to significant transformations. Our growing global network is dedicated to inventing, developing, and commercializing medicines that change lives for those with serious diseases.


At Regeneron Digital & Technology, we are expanding our AI and Platform Engineering capabilities to support next-generation intelligent systems, machine learning platforms, and cloud-native technologies. We are seeking an AI-MLOps SRE Lead Engineer to drive reliability, scalability, observability, and operational excellence across our AI, ML, and cloud ecosystem. This role will lead the design and operation of resilient platforms supporting machine learning workloads, LLMs, AI Agents, and enterprise-scale automation while enabling engineering teams to innovate with speed and confidence.


When & Where

Hyderabad (Hybrid)


Discover your role

  • Drive service reliability, availability, and performance across multi-cloud environments by establishing SLOs, SLIs, error budgets, and reliability standard methodologies.
  • Design, build, and scale enterprise ML platform infrastructure using technologies such as Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI.
  • Develop AI-powered observability solutions using anomaly detection, predictive analytics, and automated remediation to proactively identify and resolve operational issues.
  • Lead the implementation, evaluation, and monitoring of LLMs, SLMs, RAG pipelines, and AI Agent platforms, ensuring performance, governance, scalability, and operational efficiency.
  • Design and implement Infrastructure as Code (IaC), CI/CD pipelines, self-healing systems, and automation capabilities that improve engineering productivity and platform resilience.
  • Architect enterprise ChatOps solutions integrating AI/ML operations, observability platforms, operational events, and automated remediation workflows.
  • Partner with Data Science, AI Engineering, and Platform teams to deliver secure, scalable, and production-ready AI/ML platforms and services.
  • Evaluate emerging AI-native operational technologies and AI-assisted engineering solutions to improve platform reliability and operational excellence.
  • Conduct technical debt assessments, identify architectural risks, and provide strategic recommendations to modernize enterprise platforms.
  • Serve as a technical leader and trusted advisor, mentoring engineers and influencing SRE, MLOps, cloud platform, and AI operations strategy across the organization.

This role requires

  • Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence, or a related field; Master's degree preferred.
  • 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology subject areas within enterprise-scale environments.
  • Strong hands-on experience working across two or more cloud platforms, including AWS, GCP, and Azure.
  • Deep expertise with ML platform technologies such as Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI.
  • Proven experience implementing end-to-end ML workflows, including model training, deployment, experiment tracking, monitoring, and pipeline orchestration.
  • Hands-on experience applying ML techniques such as anomaly detection, predictive analytics, time-series modeling, and operational intelligence within enterprise platforms.
  • Advanced proficiency with Infrastructure as Code tools, including Terraform, Pulumi, AWS CDK, and modern CI/CD automation practices.
  • Strong programming and scripting capabilities in Python, Go, Bash, or similar languages.
  • Experience building enterprise observability platforms using Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing, metrics, and logging solutions.
  • Proven expertise designing and implementing enterprise ChatOps architectures, operational workflows, and AI-enabled automation solutions.
  • Strong ability to assess technical debt, influence technical strategy, and drive platform modernization initiatives.
  • Experience using AI tools, LLM-powered assistants, and AI Agents to enhance engineering productivity and platform operations.
  • Experience with Kubernetes and container orchestration technologies such as EKS, GKE, or AKS preferred.
  • Familiarity with MLOps and LLM evaluation technologies including Kubeflow, Feast, LangSmith, RAGAS, Evidently AI, MLflow, and Weights & Biases preferred.
  • Knowledge of cloud cost optimization, FinOps, policy-as-code, compliance automation, and multi-cloud governance practices preferred.


Does this sound like you? Apply now to take your first step towards living the Regeneron Way! We are committed to building a workplace with an inclusive culture. Regeneron is an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion or belief (or lack thereof), sex, sexual orientation, gender identity or expression, gender reassignment, marital or civil partnership status, civil status, pregnancy or parental status, age, disability, nationality, citizenship status, ethnic or national origin, membership of the Traveler community, familial status, genetic information, military or veteran status, or any other characteristic protected under applicable law. Where required, we will provide reasonable accommodation to applicants with known disabilities or chronic illnesses during the recruitment process, unless such accommodation would impose undue hardship.

 

Where necessary, we disclose salary ranges for roles in all countries in which we operate. The final offer will be determined within the relevant range based on the country of employment, specific role level, and your skills and experience. In some countries, collective bargaining agreements (CBAs) may apply and influence certain elements of pay or benefits. Regeneron offers a competitive and comprehensive total rewards package which may include, depending on country and role: annual bonuses or other incentive plans, equity awards, pension or retirement benefits, 401(k) company match, health and wellness programs, fitness centers, insurance benefits (e.g. medical, dental, vision, life and disability), paid time off, and family support benefits. For additional information about Regeneron benefits in the U.S., please visit https://careers.regeneron.com/en/working-at-regeneron/total-rewards/. For other locations, additional information will be provided during the recruitment process. If you have any questions, please speak with your recruiter. 


Please be advised that at Regeneron, we believe we do our best work when we are together. For that reason, many roles are required to be performed on‑site. Please speak with your recruiter and hiring manager for more information about on‑site expectations for your role and location.


As part of the recruitment process, certain background checks may be conducted in accordance with the laws of the country where the position is based. The purpose of such checks is to verify certain information prior to the commencement of employment such as identity, right to work and educational qualifications.


For jobs in Canada: this posting is for an existing position.

Get AI/MLOps SRE Lead Engineer jobs like this→

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Toss logo

Toss

New

Data Engineer

Seoul
✓ From careers page· 8h ago
Peloton logo

Senior Software Engineer

$173k–$214kNew York, NY
✓ From careers page· 8h ago
1GLOBAL logo

Senior Site Reliability Engineer

Berlin, Berlin
✓ From careers page· 9h ago
FactSet logo

Principal Site Reliability Engineer (Remote)

$190k–$220kNorwalk, CT
✓ From careers page· 9h ago

Frequently asked questions

What skills are required for AI/MLOps SRE Lead Engineer at Regeneron?

The required skills for AI/MLOps SRE Lead Engineer at Regeneron include: SRE, DevOps, Machine Learning, LLM, Databricks, CI/CD, Python, Go, Bash, Prometheus, Grafana, Datadog, OpenTelemetry, Kubernetes, MLOps, MLflow.

What is the seniority level for AI/MLOps SRE Lead Engineer at Regeneron?

AI/MLOps SRE Lead Engineer at Regeneron is a Lead / Senior level position.

How do I apply for AI/MLOps SRE Lead Engineer at Regeneron?

You can view the full description and apply for AI/MLOps SRE Lead Engineer at Regeneron on EchoJobs: https://echojobs.io/job/regeneron-ai-mlops-sre-lead-engineer-mk566.