DTCC logo

Director of IAM Site Reliability Engineering & Observability

DTCC

Hybrid
London, United Kingdom
Full-time
Director
Manager
10+ yrs
Salary not listedPosted 4h ago

Real job — pulled straight from DTCC’s careers page · Verified August 13, 2026 · No reposts.

Job description

DTCC is hiring a Director of IAM Site Reliability Engineering & Observability — a full-time, based in London, United Kingdom role. Apply directly on DTCC's careers page below.

IAM Engineering Director

Location: LONDON, United Kingdom

Are you ready to make an impact at DTCC?  

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We're committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

Pay and Benefits:

  • Competitive compensation, including base pay and annual incentive
  • Comprehensive health and life insurance and well-being benefits
  • Pension
  • Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
  • DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (Onsite Tuesdays, Wednesdays and a third day of your choosing)

The impact you will have in this role:

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.

Role Summary:

We are seeking an experienced Director of IAM Site Reliability Engineering (SRE) & Observability to lead the reliability, availability, operational architecture, monitoring strategy, and resiliency of the enterprise Identity and Access Management (IAM) ecosystem.

This leader will be responsible for ensuring the health, performance, scalability, recoverability, and operational readiness of critical IAM platforms including Privileged Access Management (PAM), Active Directory, Certificate Management Infrastructure (PKI), Secrets Management, Authentication Services, Cloud Identity Platforms, and related Identity Security services.

The role will establish enterprise-wide observability standards, develop service reliability architectures, define monitoring frameworks, and partner closely with Engineering teams to ensure IAM platforms are designed, instrumented, and operated for maximum availability and resilience.

The successful candidate will drive a proactive reliability culture focused on monitoring, telemetry, automation, service health, incident prevention, operational architecture, and continuous improvement.

Your Primary Responsibilities:

IAM Site Reliability Engineering Leadership

  • Lead the IAM Site Reliability Engineering (SRE) function across all IAM platforms and services.
  • Own platform availability, service health, resiliency, observability, and operational readiness objectives.
  • Establish reliability engineering practices and operational excellence standards across the IAM ecosystem.
  • Partner with engineering teams throughout the software and platform lifecycle to embed reliability-by-design principles.

Observability & Monitoring Strategy

  • Define and implement enterprise observability standards across all IAM platforms.
  • Develop monitoring architecture standards covering:
    • Infrastructure Monitoring
    • Application Monitoring
    • User Experience Monitoring
    • Transaction Monitoring
    • Dependency Monitoring
    • Security Event Monitoring
    • Cloud Service Monitoring
  • Establish standards for logging, metrics, tracing, dashboards, alerting, correlation, and telemetry collection.
  • Drive implementation of centralized observability platforms leveraging tools such as Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, App Insights, or equivalent solutions.
  • Define platform-specific monitoring requirements and operational instrumentation standards for all IAM services.
  • Ensure all IAM applications meet monitoring, alerting, and observability requirements prior to production deployment.

Reliability Architecture & Engineering

  • Create operational architecture diagrams, service dependency maps, data flow diagrams, and platform resiliency models for IAM services.
  • Work closely with IAM Engineering teams to review solution designs and identify availability, scalability, and resiliency risks.
  • Establish architecture standards for:
    • High Availability (HA)
    • Disaster Recovery (DR)
    • Multi-site Resiliency
    • Failover Design
    • Capacity Planning
    • Fault Tolerance
    • Service Recovery
  • Conduct reliability design reviews and production readiness assessments for IAM platforms and applications.

Availability & Resilience Management

  • Establish and govern Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and Availability Targets.
  • Drive initiatives to improve platform uptime, stability, recoverability, and performance.
  • Identify and eliminate single points of failure and operational bottlenecks.
  • Lead resiliency testing exercises, failover validation, disaster recovery testing, and business continuity preparedness.

Incident & Service Health Management

  • Lead major incident management, executive communications, and post-incident reviews.
  • Analyze incident trends, recurring failures, and reliability gaps to drive platform improvements.
  • Establish proactive service health review processes and operational risk assessments.
  • Ensure corrective actions are tracked and implemented to prevent repeat incidents.

Automation & Operational Excellence

  • Drive automation of monitoring, alert triage, remediation, service validation, and operational workflows.
  • Promote Infrastructure as Code (IaC), automated health checks, self-healing capabilities, and operational engineering practices.
  • Improve alert quality by reducing noise and focusing on actionable service indicators.
  • Establish reliability scorecards and operational maturity frameworks across IAM platforms.

Cross-Functional Leadership

  • Partner with IAM Engineering, Security Engineering, Enterprise Architecture, Cloud, Infrastructure, Network, and Vendor teams.
  • Serve as the primary authority for IAM platform observability, availability, and operational architecture standards.
  • Influence engineering roadmaps by incorporating reliability, monitoring, and resiliency requirements into platform design and development processes.

**NOTE: The Primary Responsibilities of this role are not limited to the details above. **

Qualifications:

  • Bachelor's degree preferred or equivalent experience

Talents Needed For Success:

  • Minimum of 10 years related experience
  • 12+ years of experience in Site Reliability Engineering, Platform Engineering, IAM, Infrastructure Engineering, or Cybersecurity.
  • Experience supporting complex IAM environments including PAM, Active Directory, PKI, Authentication Services, Secrets Management, and Cloud Identity platforms.
  • Strong expertise in observability, monitoring architecture, telemetry design, and platform instrumentation.
  • Experience creating architecture diagrams, dependency maps, operational blueprints, and resiliency designs.
  • Hands-on experience with enterprise monitoring and observability platforms.
  • Strong knowledge of High Availability, Disaster Recovery, Business Continuity, and resilience engineering.
  • Experience leading major incident management and operational transformation programs.
  • Proven ability to influence architecture and engineering teams without direct ownership.
  • Strong executive communication and stakeholder management skills.

We offer top class training and development for you to be an asset in our organization!

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

 

We are seeking an experienced Director of IAM Site Reliability Engineering (SRE) & Observability to lead the reliability, availability, operational architecture, monitoring strategy, and resiliency of the enterprise Identity and Access Management (IAM) ecosystem.

Get Director of IAM Site Reliability Engineering & Observability jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Walmart logo

Staff Data Scientist

$110k–$220kBentonville, AR
✓ From careers page· 17m ago
Westernalliancebank logo

Principal Engineer, AIOps

Phoenix, AZ
✓ From careers page· 46m ago
Nike logo

Nike

New

Lead Software Engineer

Karnataka, India
✓ From careers page· 46m ago
NIX logo

NIX

New

Lead DevOps Engineer

Poland
✓ From careers page· 47m ago

Frequently asked questions

What skills are required for Director of IAM Site Reliability Engineering & Observability at DTCC?

The required skills for Director of IAM Site Reliability Engineering & Observability at DTCC include: SRE, IAM, Splunk, Datadog, Grafana, Active Directory.

What is the seniority level for Director of IAM Site Reliability Engineering & Observability at DTCC?

Director of IAM Site Reliability Engineering & Observability at DTCC is a Director / Manager level position.

How do I apply for Director of IAM Site Reliability Engineering & Observability at DTCC?

You can view the full description and apply for Director of IAM Site Reliability Engineering & Observability at DTCC on EchoJobs: https://echojobs.io/job/dtcc-iam-engineering-director-110tn.