
Senior Principal Cloud Development Engineer
Real job — pulled straight from DTCC’s careers page · Verified July 18, 2026 · No reposts.
Job description
DTCC is hiring a Senior Principal Cloud Development Engineer — a full-time, based in London, United Kingdom role. Apply directly on DTCC's careers page below.
Senior Principal Cloud Development Engineer
Location: LONDON, United Kingdom
Are you ready to make an impact at DTCC?
Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We're committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.
The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.
Pay and Benefits:
- Competitive compensation, including base pay and annual incentive
- Comprehensive health and life insurance and well-being benefits
- Pension
- Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
- DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).
The impact you will have in this role
The Senior Principal, Operations Resiliency & GameDay Strategy within Cloud Operations is responsible for defining and operating the enterprise GameDay program that validates the resiliency of cloud-hosted applications and platforms. You will design failure scenarios drawn from real production incidents, select and prioritize applications for testing, and ensure that exercise outcomes translate directly into improved runbooks, recovery paths, and architectural resilience. Working across Cloud Operations, Incident Management, and application engineering teams, you will drive a closed-loop process where what breaks in production is systematically tested, validated, and retested until recovery capabilities are proven and measurable.
Your primary responsibilities
- Define and operationalize the enterprise GameDay strategy that validates resiliency of cloud-hosted applications and platforms.
- Establish application selection criteria, testing frequency, and a scalable operating model for controlled failure testing.
- Develop a taxonomy of application patterns and map them to targeted, reusable fault injection scenarios drawn from real production failures.
- Oversee GameDay execution, manual and automated, validating observability, detection, response, and recovery workflows.
- Close the loop between production incidents and simulated testing: incident, scenario, test, finding, mitigation. retest.
- Drive identification of missing recovery paths, observability gaps, and ineffective automation, then ensure findings are tracked and assigned to accountable owners.
- Push execution from manual to automated: fault injection, scenario orchestration, metrics capture, and reporting.
- Define and track effectiveness metrics: time to detect, time to recover, automated recovery success rates, and reduction in repeat incident patterns.
- Partner across Incident Management, Application Teams, Platform Engineering, and SRE to ensure GameDay outputs translate directly into improved runbooks, recovery paths, and architectural resilience.
Qualifications:
- Minimum of 10 years of related experience
- Bachelor's degree preferred or equivalent experience
Talents Needed for Success:
- Deep experience in cloud operations, SRE, or resiliency engineering
- Hands-on knowledge of incident management and root cause analysis
- Experience with fault injection tools such as AWS FIS, Gremlin, or equivalent
- Experience with AI-assisted operations tooling such as AWS DevOps Agent or equivalent
- Familiarity with policy-as-code frameworks such as OPA, Sentinel, or equivalent
- Strong understanding of distributed systems failure modes, observability architectures, and automation/runbook engineering
- Ability to operate horizontally across engineering and operations organizations — influencing without direct delivery ownership
- Strong judgment on what to test, when, and why
We offer top class training and development for you to be an asset in our organization!
We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.
About Us
DTCC proudly supports Flexible Work Arrangements favoring openness and gives people freedom to do their jobs well, by encouraging diverse opinions and emphasizing teamwork. When you join our team, you’ll have an opportunity to make meaningful contributions at a company that is recognized as a thought leader in both the financial services and technology industries. A DTCC career is more than a good way to earn a living. It’s the chance to make a difference at a company that’s truly one of a kind.
Learn more about Clearance and Settlement by clicking here.
About the Organization
Serves as a dedicated technology resource for advancing DTCC’s business opportunities and providing industry thought leadership for leveraging new technology. The goal of this new department is to partner internally with IT, our business and regulatory divisions and externally with clients, regulators, and fintech vendors, to help build new platforms and business models to advance DTCC’s mission to support the financial markets.
The Senior Principal, Operations Resiliency & GameDay Strategy within Cloud Operations is responsible for defining and operating the enterprise GameDay program that validates the resiliency of cloud-hosted applications and platforms. You will design failure scenarios drawn from real production incidents, select and prioritize applications for testing, and ensure that exercise outcomes translate directly into improved runbooks, recovery paths, and architectural resilience. Working across Cloud Operations, Incident Management, and application engineering teams, you will drive a closed-loop process where what breaks in production is systematically tested, validated, and retested until recovery capabilities are proven and measurable.Get Senior Principal Cloud Development Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs

Manager, Production Operations and Site Reliability Engineering



Frequently asked questions
What skills are required for Senior Principal Cloud Development Engineer at DTCC?
The required skills for Senior Principal Cloud Development Engineer at DTCC include: SRE.
What is the seniority level for Senior Principal Cloud Development Engineer at DTCC?
Senior Principal Cloud Development Engineer at DTCC is a Senior / Principal level position.
How do I apply for Senior Principal Cloud Development Engineer at DTCC?
You can view the full description and apply for Senior Principal Cloud Development Engineer at DTCC on EchoJobs: https://echojobs.io/job/dtcc-senior-principal-cloud-development-engineer-54tai.