DTCC logo

Senior Principal Cloud Development Engineer

DTCC

Hybrid
London, United Kingdom
Full-time
Senior
Principal
10+ yrs
Salary not listedPosted 4w ago

Real job — pulled straight from DTCC’s careers page · Verified July 18, 2026 · No reposts.

Job description

DTCC is hiring a Senior Principal Cloud Development Engineer — a full-time, based in London, United Kingdom role. Apply directly on DTCC's careers page below.

Senior Principal Cloud Development Engineer

Location: LONDON, United Kingdom

Are you ready to make an impact at DTCC?

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We're committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.

Pay and Benefits:

  • Competitive compensation, including base pay and annual incentive
  • Comprehensive health and life insurance and well-being benefits
  • Pension 
  • Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
  • DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee). 

The impact you will have in this role

The Senior Principal, Operations Resiliency & GameDay Strategy within Cloud Operations is responsible for defining and operating the enterprise GameDay program that validates the resiliency of cloud-hosted applications and platforms. You will design failure scenarios drawn from real production incidents, select and prioritize applications for testing, and ensure that exercise outcomes translate directly into improved runbooks, recovery paths, and architectural resilience. Working across Cloud Operations, Incident Management, and application engineering teams, you will drive a closed-loop process where what breaks in production is systematically tested, validated, and retested until recovery capabilities are proven and measurable.

Your primary responsibilities

  • Define and operationalize the enterprise GameDay strategy that validates resiliency of cloud-hosted applications and platforms.
  • Establish application selection criteria, testing frequency, and a scalable operating model for controlled failure testing.
  • Develop a taxonomy of application patterns and map them to targeted, reusable fault injection scenarios drawn from real production failures.
  • Oversee GameDay execution, manual and automated, validating observability, detection, response, and recovery workflows.
  • Close the loop between production incidents and simulated testing: incident, scenario, test, finding, mitigation. retest.
  • Drive identification of missing recovery paths, observability gaps, and ineffective automation, then ensure findings are tracked and assigned to accountable owners.
  • Push execution from manual to automated: fault injection, scenario orchestration, metrics capture, and reporting.
  • Define and track effectiveness metrics: time to detect, time to recover, automated recovery success rates, and reduction in repeat incident patterns.
  • Partner across Incident Management, Application Teams, Platform Engineering, and SRE to ensure GameDay outputs translate directly into improved runbooks, recovery paths, and architectural resilience.

 

Qualifications:

  • Minimum of 10 years of related experience
  • Bachelor's degree preferred or equivalent experience

Talents Needed for Success:

  • Deep experience in cloud operations, SRE, or resiliency engineering
  • Hands-on knowledge of incident management and root cause analysis
  • Experience with fault injection tools such as AWS FIS, Gremlin, or equivalent
  • Experience with AI-assisted operations tooling such as AWS DevOps Agent or equivalent
  • Familiarity with policy-as-code frameworks such as OPA, Sentinel, or equivalent
  • Strong understanding of distributed systems failure modes, observability architectures, and automation/runbook engineering
  • Ability to operate horizontally across engineering and operations organizations — influencing without direct delivery ownership
  • Strong judgment on what to test, when, and why

We offer top class training and development for you to be an asset in our organization!

We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

About Us

With over 50 years of experience, DTCC is the premier post-trade market infrastructure for the global financial services industry. From 20 locations around the world, DTCC, through its subsidiaries, automates, centralizes, and standardizes the processing of financial transactions, mitigating risk, increasing transparency, enhancing performance and driving efficiency for thousands of broker/dealers, custodian banks and asset managers. Industry owned and governed, the firm innovates purposefully, simplifying the complexities of clearing, settlement, asset servicing, transaction processing, trade reporting and data services across asset classes, bringing enhanced resilience and soundness to existing financial markets while advancing the digital asset ecosystem. In 2024, DTCC’s subsidiaries processed securities transactions valued at U.S. $3.7 quadrillion and its depository subsidiary provided custody and asset servicing for securities issues from over 150 countries and territories valued at U.S. $99 trillion. DTCC’s Global Trade Repository service, through locally registered, licensed, or approved trade repositories, processes more than 25 billion messages annually. To learn more, please visit us at www.dtcc.com or connect with us on LinkedIn, X, YouTube, Facebook and Instagram.

DTCC proudly supports Flexible Work Arrangements favoring openness and gives people freedom to do their jobs well, by encouraging diverse opinions and emphasizing teamwork. When you join our team, you’ll have an opportunity to make meaningful contributions at a company that is recognized as a thought leader in both the financial services and technology industries. A DTCC career is more than a good way to earn a living. It’s the chance to make a difference at a company that’s truly one of a kind.

Learn more about Clearance and Settlement by clicking here.

About the Organization

Serves as a dedicated technology resource for advancing DTCC’s business opportunities and providing industry thought leadership for leveraging new technology. The goal of this new department is to partner internally with IT, our business and regulatory divisions and externally with clients, regulators, and fintech vendors, to help build new platforms and business models to advance DTCC’s mission to support the financial markets.

The Senior Principal, Operations Resiliency & GameDay Strategy within Cloud Operations is responsible for defining and operating the enterprise GameDay program that validates the resiliency of cloud-hosted applications and platforms. You will design failure scenarios drawn from real production incidents, select and prioritize applications for testing, and ensure that exercise outcomes translate directly into improved runbooks, recovery paths, and architectural resilience. Working across Cloud Operations, Incident Management, and application engineering teams, you will drive a closed-loop process where what breaks in production is systematically tested, validated, and retested until recovery capabilities are proven and measurable.

Get Senior Principal Cloud Development Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Alcon logo

Manager, Production Operations and Site Reliability Engineering

$140k–$182kLake Forest, CA
✓ From careers page· 13m ago
Moniepoint logo

Senior Site Reliability Engineer

Remote · Nigeria-eligible
✓ From careers page· 17m ago
Encora logo

Senior Java Developer

Lima, PE
✓ From careers page· 17m ago
Encora logo

Senior Java Developer

Lima, Peru
✓ From careers page· 17m ago

Frequently asked questions

What skills are required for Senior Principal Cloud Development Engineer at DTCC?

The required skills for Senior Principal Cloud Development Engineer at DTCC include: SRE.

What is the seniority level for Senior Principal Cloud Development Engineer at DTCC?

Senior Principal Cloud Development Engineer at DTCC is a Senior / Principal level position.

How do I apply for Senior Principal Cloud Development Engineer at DTCC?

You can view the full description and apply for Senior Principal Cloud Development Engineer at DTCC on EchoJobs: https://echojobs.io/job/dtcc-senior-principal-cloud-development-engineer-54tai.