Resideo logo

Director, Site Reliability Engineering & Cloud Operations

Resideo

Hybrid
Austin, TX
Full-time
Manager
Principal
15+ yrs
Salary not listedPosted 1mo ago

Real job — pulled straight from Resideo’s careers page · Verified July 12, 2026 · No reposts.

Job description

Resideo is hiring a Director, Site Reliability Engineering & Cloud Operations — a full-time, based in Austin, TX role. Apply directly on Resideo's careers page below.

Director, Site Reliability Engineering & Cloud Operations (SRE)

Location: Austin, TX, United States; Hybrid

At Resideo, we imagine a world where homes and buildings are good for the planet, and where technology works to simplify everyday life. In that world, people are healthy, happy, and secure. To help create this future, we will work every day to simplify the connected world so people have peace of mind and can focus on what matters most. Resideo is making a large investment in our engineering group. With global reach and impact, we are dedicated to an investment in building our team as we develop new products and introduce them to consumers around the world (NPI). Being an established leader in the connected products space, we will give you a platform to work on new and innovative projects as a member of a team of intelligent innovators that are developing products that truly align with our mission of protecting what matters most.

This is an exciting opportunity to lead cloud operations for one of the largest IoT ecosystems in the world, shaping the future of cloud infrastructure, SRE, and AI-driven operations. You'll work alongside world-class engineering talent and cutting-edge technologies to ensure Resideo’s mission of simplifying everyday life through innovative connected products. As a leader, you will have the opportunity to lead the platform engineering transformation in a global organization of multiple teams in delivering on business priorities while collaborating with development leaders and executives to define and advance best practices. 

Resideo is seeking a strategic and experienced leader to oversee the global cloud infrastructure, Site Reliability Engineering (SRE) for our large-scale, connected products ecosystem and CloudOps. This role will drive the performance, reliability, security, and operational excellence of our multi-cloud environments (Azure), supporting millions of IoT devices and trillions of data points and events. The ideal candidate will have deep expertise in cloud infrastructure, IoT, and large-scale SaaS platforms, and be passionate about fostering a culture of innovation, reliability, and automation. 

 

JOB DUTIES:

  • Cloud Infrastructure & SRE Strategy
    • Define and execute global cloud operations and SRE strategies, ensuring 99.999%+ uptime for mission-critical IoT services.
    • Architect, implement, and optimize multi-cloud infrastructure to support IoT devices with low-latency data processing, scalability, and high availability.
    • Drive cost optimization strategies while balancing performance, redundancy, and financial efficiency across cloud platforms (Azure). 
    • Develop automated deployment, monitoring, and recovery systems using technologies like Kubernetes, Terraform, Ansible, and CI/CD pipelines.
  • Reliability, Performance & Incident Management
    • Establish and refine SLOs, SLIs, and KPIs for service reliability, performance, and capacity planning.
    • Build and optimize incident management, disaster recovery, and resilience engineering frameworks.
    • Leverage AI/ML-driven automation for proactive failure detection and remediation.
  • Security & Compliance
    • Implement robust security practices and ensure cloud security, compliance with standards such as SOC2, GDPR, and NIST, and oversee the zero-trust security model for IoT data protection.
    • Collaborate with security and compliance teams to manage risk and ensure regulatory adherence across cloud platforms.
  • Team Leadership & Cross-Functional Collaboration
    • Lead and mentor a global team of Cloud Engineers, SREs, and SW professionals, fostering a culture of continuous learning and innovation.
    • Partner with product management, software engineering, and customer support to optimize IoT device onboarding, firmware updates, and cloud-to-edge performance.
    • Collaborate with finance and executive leadership to develop long-term cloud investment strategies.

 

YOU MUST HAVE:

  • 15 + years in Computer Science, Electrical Engineering, or a related field
  • 15+ years of experience in Cloud Operations, SRE, or Infrastructure Engineering, with 8+ years in technical leadership roles
  • 5+ years of experience managing large-scale, distributed IoT cloud environments supporting billions of data points per day
  • 5+ years of deep professional experience in Azure cloud platforms including networking, storage, compute, and database services
  • 5+ years of experience in Kubernetes, Terraform, CI/CD pipelines, and observability tools (e.g., Prometheus, Grafana, ELK, etc.)
  • 5+ years of experience in large-scale systems design and architecture, with a focus on reliability, performance, and scalability of cloud-native platforms
  • 5+ years of hands-on experience with tools like Terraform, Ansible, CDK, Pulumi for Infrastructure-as-Code (IaC), and managing cloud-native architectures

 

WE VALUE:

  • Strong background in AI/ML-driven automation for cloud infrastructure monitoring, self-healing, and optimization
  • Solid understanding of security-first cloud architectures, DevSecOps, and compliance standards (SOC2, GDPR, NIST)
  • Proven ability to manage teams across multiple global time zones, ensuring operational excellence and driving performance in large, distributed environments
  • Expertise in incident management, disaster recovery, and building resilience engineering frameworks
  • Ability and desire to review code, system designs, and engage in system engineering discussions and decisions
  • Experience managing Consumer IoT ecosystems with large-scale sensor data processing and real-time analytics
  • Expertise in serverless architecture, edge computing, and IoT protocol optimization
  • Strong financial acumen in cloud cost management, and forecasting
  • Familiarity with regulatory compliance frameworks such as SOC2, GDPR, and ISO 27001
  • Relevant certifications, such as Azure Expert

 

WHAT'S IN IT FOR YOU:

  • Innovation: Bring your creative ideas to the table and be part of a company that values out-of-the-box thinking

 

#LI-HYBRID

#LI-MA1 

About Us

Resideo is a global leader in smart home and building solutions, with trusted brands including Honeywell Home, First Alert, and Resideo helping people feel more comfortable, secure, connected, and in control every day. Our products and technologies are found in more than 150 million homes and businesses worldwide. From intelligent climate solutions to security, sensing, water, and connected home technologies, Resideo develops and manufactures products designed to simplify everyday life and help protect what matters most. Our global teams span engineering, manufacturing, software, product management, supply chain, customer experience, and more — all working together to shape the future of connected living through innovation, quality, and meaningful real-world impact. At Resideo, our teams help create products and experiences that make everyday life more comfortable, secure, and connected for millions around the world. Learn more at www.resideo.com.

You can find out more about how the talent community works here: Resideo Talent Community Terms. Our recruitment privacy notice Resideo -Recruitment Privacy Notice - Dec 16 2022 describes in more detail how we process your personal data and how you can exercise your personal data rights.

If a disability prevents you from applying for a job through our website, request assistance here.

Resideo is seeking a strategic and experienced leader to oversee the global cloud infrastructure, Site Reliability Engineering (SRE) for our large-scale, connected products ecosystem and CloudOps. This role will drive the performance, reliability, security, and operational excellence of our multi-cloud environments (Azure), supporting millions of IoT devices and trillions of data points and events. The ideal candidate will have deep expertise in cloud infrastructure, IoT, and large-scale SaaS platforms, and be passionate about fostering a culture of innovation, reliability, and automation.

Get Director, Site Reliability Engineering & Cloud Operations jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
AbbVie logo

Data Engineer

North Chicago, IL
✓ From careers page· 30m ago
Thomson Reuters logo

Senior Cloud Security Engineer

Bengaluru, Karnataka
✓ From careers page· 31m ago
Waystar logo

Senior Data Scientist

Lehi, UT
✓ From careers page· 2h ago
TetraScience logo

Scientific Data Architect

Frankfurt am Main, Hessen, Germany
✓ From careers page· 2h ago

Frequently asked questions

What skills are required for Director, Site Reliability Engineering & Cloud Operations at Resideo?

The required skills for Director, Site Reliability Engineering & Cloud Operations at Resideo include: Azure, Kubernetes, Terraform, Ansible, CI/CD, Prometheus, Grafana, Elasticsearch, SOC 2, GDPR, NIST, IoT, SRE.

What is the seniority level for Director, Site Reliability Engineering & Cloud Operations at Resideo?

Director, Site Reliability Engineering & Cloud Operations at Resideo is a Manager / Principal level position.

How do I apply for Director, Site Reliability Engineering & Cloud Operations at Resideo?

You can view the full description and apply for Director, Site Reliability Engineering & Cloud Operations at Resideo on EchoJobs: https://echojobs.io/job/resideo-director-site-reliability-engineering-cloud-operations-sre-pp2u0.