Wells Fargo logo

Senior Manager, Site Reliability Engineering

Wells Fargo

On-site
Irving, TX
Full-time
Manager
Senior
7+ yrs
Salary not listedPosted 5d ago

Real job — pulled straight from Wells Fargo’s careers page · Verified August 8, 2026 · No reposts.

Job description

Wells Fargo is hiring a Senior Manager, Site Reliability Engineering — a full-time, based in Irving, TX role. Apply directly on Wells Fargo's careers page below.

Senior Manager, Site Reliability Engineering

Location: IRVING, TX

Time Type: Full time

Job Description

About this role:

Wells Fargo is seeking a Systems Operations Senior Manager, Site Reliability Engineering (SRE) to lead a team of engineers responsible for the reliability, availability, performance, scalability, and operational excellence of critical customer-facing and enterprise technology platforms. This leader drives resiliency engineering, observability, incident management, automation, and continuous service improvement while partnering closely with Application Development, Infrastructure, Cybersecurity, Product, and Business teams.

The role combines deep technical expertise with strong leadership to establish reliability engineering practices, reduce operational risk, improve customer experience, and accelerate delivery through automation and engineering excellence.


In this role, you will:

Reliability & Platform Stability

  • Own the reliability strategy for critical applications and platforms.

  • Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.

  • Drive platform availability, resiliency, recoverability, and scalability initiatives.

  • Reduce Mean Time to Detect (MTTD) and Mean Time to Restore (MTTR) through proactive engineering.

  • Establish resiliency standards, operational readiness reviews, and production certification requirements.

Incident Management & Problem Management

  • Lead major incident response activities for high-severity customer-impacting events.

  • Drive root cause analysis (RCA) and corrective action management.

  • Establish operational governance and escalation processes.

  • Identify systemic reliability risks and remediation opportunities.

  • Coordinate cross-functional recovery efforts involving internal teams and third-party vendors.

Observability & Monitoring

  • Define enterprise observability standards and best practices.

  • Drive implementation of:

    • Metrics

    • Logging

    • Tracing

    • Synthetic Monitoring

    • Business Observability

  • Develop operational dashboards and executive-level service health reporting.

  • Ensure end-to-end visibility across customer journeys and critical business processes.

Automation & Engineering Excellence

  • Lead initiatives to automate operational processes and reduce manual effort.

  • Drive adoption of Infrastructure as Code (IaC), CI/CD, Auto-Remediation, and AIOps capabilities.

  • Improve operational efficiency through self-healing platforms and intelligent alerting.

  • Eliminate repetitive operational tasks through engineering solutions.

Vendor & Third-Party Reliability Management

  • Establish operational engagement standards for strategic vendors and partners.

  • Drive vendor accountability for production stability, incident response, and RCA delivery.

  • Participate in contractual resiliency reviews and SLA governance.

  • Ensure third-party technology providers meet operational and reliability expectations.

Risk & Compliance

  • Partner with Risk, Audit, Cybersecurity, and Compliance teams.

  • Ensure platforms meet regulatory, resiliency, and operational risk requirements.

  • Maintain production support controls, procedures, and evidence for audits and examinations.


Required Qualifications:

  • 7+ years of Systems Engineering and Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education

  • 3+ years of management or leadership experience

  • 5+ years of leadership experience managing engineering or SRE teams

  • 5 years experience supporting large-scale, customer-facing applications and platforms

  • 5 years experience displaying strong understanding of: Incident Management, Problem Management, SRE Principles, DevOps Practices, Cloud Technologies, Production Operations

Desired Qualifications:

  • Financial services or highly regulated industry experience.

  • Expertise with: Splunk, Grafana, AppDynamics, Dynatrace, OpenTelemetry, Prometheus, Kubernetes/OpenShift, Public Cloud Platforms (AWS, Azure, GCP)

  • Experience implementing Business Observability and AIOps solutions.

  • Knowledge of ITIL, Operational Resiliency, Disaster Recovery, and Capacity Management frameworks.


Locations:

  • 794 Davis St, San Leandro, California

  • 401 Las Colinas Blvd W. Irving, Texas

Posting End Date: 

12 Aug 2026

*Job posting may come down early due to volume of applicants.

We Value Equal Opportunity

Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.

Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit’s risk appetite and all risk and compliance program requirements.

Candidates applying to job openings posted in Canada: Applications for employment are encouraged from all qualified candidates, including women, persons with disabilities, aboriginal peoples and visible minorities. Accommodation for applicants with disabilities is available upon request in connection with the recruitment process.

Applicants with Disabilities

To request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo.

Drug and Alcohol Policy

 

Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.

Wells Fargo Recruitment and Hiring Requirements:

a. Third-Party recordings are prohibited unless authorized by Wells Fargo.

b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.

Get Senior Manager, Site Reliability Engineering jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
ThoughtSpot logo

Senior Systems Reliability Engineer

Bengaluru, India
✓ From careers page· 2h ago
JLL logo

JLL

New

Senior Cloud Architect & DevOps Manager

$212k–$259kRemote · US-eligible
✓ From careers page· 3h ago
Gartner logo

Principal Site Reliability Engineer

Chennai
✓ From careers page· 5h ago

Frequently asked questions

What skills are required for Senior Manager, Site Reliability Engineering at Wells Fargo?

The required skills for Senior Manager, Site Reliability Engineering at Wells Fargo include: SRE, Splunk, Grafana, OpenTelemetry, Prometheus, Kubernetes, OpenShift, AWS, Azure, GCP, ITIL.

What is the seniority level for Senior Manager, Site Reliability Engineering at Wells Fargo?

Senior Manager, Site Reliability Engineering at Wells Fargo is a Manager / Senior level position.

How do I apply for Senior Manager, Site Reliability Engineering at Wells Fargo?

You can view the full description and apply for Senior Manager, Site Reliability Engineering at Wells Fargo on EchoJobs: https://echojobs.io/job/wells-fargo-senior-manager-site-reliability-engineering-5s46d.