JPMorgan Chase logo

Senior Lead Site Reliability Engineer

JPMorgan Chase

On-site
Hyderabad, Telangana, India
Full-time
Senior
Manager
5+ yrs
Salary not listedPosted 24m ago

Real job — pulled straight from JPMorgan Chase’s careers page · Verified October 10, 2026 · No reposts.

Job description

JPMorgan Chase is hiring a Senior Lead Site Reliability Engineer — a full-time, based in Hyderabad, Telangana, India role. Apply directly on JPMorgan Chase's careers page below.

Sr Lead SRE - Reliability Engineering & Problem Management

Location: Hyderabad, Telangana, India

As a Senior Lead Site Reliability Engineer at JPMorgan Chase within the Reliability Engineering & Problem Management team, you will serve as the technical authority for post-incident investigations and operational resilience. You will lead deep technical reviews, validate causal analysis, and ensure corrective actions address systemic causes. You will partner with cross-functional teams to drive measurable improvements in reliability, resilience, and operational excellence. You will help ensure incidents translate into lasting engineering improvements and reduction of repeat failures.

Job responsibilities

  • Serve as the technical authority for Problem Management-led Root Cause Analysis reviews across major incidents and service-impacting events
  • Lead deep-dive RCA challenge sessions, validating technical findings and causal chains through evidence-based analysis. Assess the quality and accuracy of root cause investigations, ensuring conclusions are technically sound and defensible. Challenge assumptions, unsupported conclusions, symptom-based findings, and ineffective corrective actions. Drive a culture of accountability focused on systemic learning and long-term reliability improvements
  • Evaluate detection gaps, monitoring effectiveness, observability shortcomings, automation opportunities, resilience weaknesses, process breakdowns, and human factors. Validate that corrective actions address the true root cause and reduce recurrence likelihood and impact. Review corrective actions for closure and effectiveness, ensuring intended reliability and operational outcomes. Provide technical challenge and independent review of vendor, third-party, and internal investigation reports
  • Use enterprise-authorized AI capabilities to accelerate reliability design and operational decisioning, validating outputs and handling operational data according to sensitivity and security requirements
  • Lead reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices, ensuring traceability, auditability, resiliency, and security controls. Define and govern Service Level Objectives, Service Level Indicators, Reliability Metrics, and Error Budgets
  • Evaluate system architecture against reliability design principles
  • Recommend resilience patterns including graceful degradation, dependency isolation, rate limiting, circuit breakers, fault tolerance, capacity management, auto-remediation, and self-healing capabilities
  • Lead reliability maturity assessments across platforms and services
  • Lead enterprise-authorized AI capabilities for RCA generation, incident analysis, log analytics, pattern discovery, problem trend analysis, and corrective action recommendations
  • Establish governance for explainability, auditability, data handling, and validation of AI-generated findings
  • Drive AI-assisted reliability workflows across the incident lifecycle

 

Required qualifications, capabilities and skills

  • Formal training or certification on security engineering concepts and 5+ years applied experience. Hands on experience in Infrastructure Engineering, Site Reliability Engineering, Production Engineering, Systems Engineering, or Software Engineering. Strong exposure to leading or supporting critical incident investigations
  • Experience conducting deep technical RCAs for enterprise-scale environments, working in highly regulated and mission-critical environments
  • Advanced expertise in Network Engineering, Cloud Infrastructure, Linux/Windows Platforms, Middleware Technologies, Database Technologies, Storage Platforms, Application Architecture, DevOps Toolchains, Distributed Systems, and Enterprise Monitoring Platforms
  • Strong hands-on expertise in SLO/SLI Engineering, Distributed Tracing, Telemetry Design, Reliability Metrics, Error Budget Management, and AIOps Platforms
  • Experience with tools such as Splunk, Dynatrace, Grafana, Datadog, Prometheus, AppDynamics, Elastic, and Open Telemetry
  • Deep understanding of Root Cause Analysis methodologies, Five Whys, Fault Tree Analysis, Event Correlation, Human Factors Analysis, Systemic Cause Analysis, Problem Management Governance, and Major Incident Management
  • Demonstrated experience using enterprise-authorized AI capabilities to improve reliability engineering workflows with strong validation habits and awareness of data sensitivity
  • Ability to set team practices for safe AI usage in operations while maintaining resiliency, security, and auditability outcomes
  • Strong executive communication skills. Ability to challenge senior engineering stakeholders constructively. Proven ability to influence without direct authority
  • Ability to translate technical findings into executive-ready narratives

 

Preferred qualifications, capabilities and skills

  • Experience leading reliability engineering initiatives in large-scale, complex environments
  • Expertise in AI-enabled incident analysis and reliability workflows
  • Experience establishing governance for explainability and auditability of AI-generated findings
  • Demonstrated ability to drive measurable improvements in reliability, resilience, and operational excellence
Lead post-incident investigations and drive reliability improvements through advanced SRE and Problem Management practices.

Get Site Reliability Engineer jobs like this→

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
ING logo

ING

New

Site Reliability Engineer

Bucharest, RO
✓ From careers page· 1h ago
NBCUniversal logo

Senior Data Engineer

$115k–$145kNew York, NY
✓ From careers page· 3h ago
Citigroup logo

Application Development Lead Analyst

Pune, Maharashtra, India
✓ From careers page· 3h ago
Citigroup logo

Public Cloud Forward Deployed Engineer

Belfast, United Kingdom
✓ From careers page· 3h ago

Frequently asked questions

What skills are required for Senior Lead Site Reliability Engineer at JPMorgan Chase?

The required skills for Senior Lead Site Reliability Engineer at JPMorgan Chase include: SRE, Linux, Windows, DevOps, Splunk, Grafana, Datadog, Prometheus, OpenTelemetry, AI.

What is the seniority level for Senior Lead Site Reliability Engineer at JPMorgan Chase?

Senior Lead Site Reliability Engineer at JPMorgan Chase is a Senior / Manager level position.

How do I apply for Senior Lead Site Reliability Engineer at JPMorgan Chase?

You can view the full description and apply for Senior Lead Site Reliability Engineer at JPMorgan Chase on EchoJobs: https://echojobs.io/job/jpmorgan-chase-sr-lead-sre-reliability-engineering-problem-management-vopnl.