Retool logo

Site Reliability Engineer

Retool

On-site
San Francisco, CA
Full-time
Mid Level
Salary not listedPosted 2w ago

Real job — pulled straight from Retool’s careers page · Verified August 2, 2026 · No reposts.

Job description

Retool is hiring a Site Reliability Engineer — a full-time, based in San Francisco, CA role. Apply directly on Retool's careers page below.

Site Reliability Engineer (SRE)

Location: San Francisco, New York

Department: Engineering

Location Type: IN_OFFICE

Employment Type: FULL_TIME

WHY WE'RE LOOKING FOR YOU:
Good software has to run where customers need it. For many of Retool's largest customers, that means running Retool in their own infrastructure, behind their own controls, with the reliability and operational clarity they would expect from any critical system.

Retool's Core Infrastructure team owns the systems that make this possible: Retool Cloud, managed single tenant environments, BYOC (bring-your-own-cloud) environments, Kubernetes and Helm deployments, Docker Compose, and the migration paths between them. It is a broad surface area, and it is one of the biggest levers we have for making Retool work for enterprise customers.

The work is not clean-room infrastructure. Customers run different clouds, different versions, different deployment models, and different levels of operational maturity. A bad upgrade experience can leave a customer many versions behind. A manual Terraform run can become the bottleneck during a launch or incident.

We are hiring SREs who want to turn that mess into leverage. You will help us reduce customer toil, automate upgrades and infrastructure changes, build reliability tooling across Retool Cloud and customer-owned environments, and make Retool easier to deploy and operate at enterprise scale. The strongest candidates are comfortable debugging Kubernetes, Terraform, AWS, Postgres, networking, and deployment problems, then stepping back and building the automation or product surface that prevents the same problem from happening again.

What you'll do:
  • Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
  • Build the automation that turns today's manual infrastructure work into repeatable systems: Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
  • Improve observability for Retool Cloud, self-hosted customers, and internal operators. We care less about exposing every metric and more about turning health signals into clear status, likely causes, and recommended actions.
  • Design safer deployment, upgrade, and rollback paths so Cloud and managed customers can stay current
  • Help move customers from legacy or less-supported deployment models toward supported paths such as Retool's official deployment paths (Blueprints, Kubernetes, and Helm), with migration flows that are repeatable enough for customers, Support, and TAMs to trust.
  • Partner with product engineers on infrastructure requirements for new Retool products, especially when they introduce new dependencies
  • Lead through ambiguity, make careful risk calls, and communicate clearly while things are moving quickly.
  • Write the docs, runbooks, design notes, and migration guides that make complex systems understandable to other engineers and to customers.

What we're looking for:
Infrastructure fundamentals
  • Deep experience operating production infrastructure in AWS.
  • Experience improving reliability for customer-facing SaaS systems.
  • Strong Kubernetes fundamentals.
  • Real Terraform or infrastructure-as-code experience.
  • Good operational judgment around databases, especially Postgres.

Reliability and automation
  • Experience building or operating observability systems.
  • Programming ability in a language such as Go, Python, TypeScript, Java, or Ruby.
  • A bias toward automation. If you find yourself doing the same operational task twice, you should start thinking about the interface, workflow, or tool that eliminates the third time.

Customer and team judgment
  • Clear written communication.
  • Comfort working directly with customer-facing teams and, when useful, customers themselves.

What makes SREs successful here:
You will do well here if you like infrastructure that sits close to real customer pain. Some days that means debugging a specific customer environment. Other days it means improving Retool Cloud reliability or designing the migration path so the next 25 customers do not need that same debugging session.

We value SREs who are ambitious, curious, energetic, and careful with the details. Retool moves quickly, priorities can change, and the systems are not always as clean as we want them to be. The work needs SREs who can get their hands dirty, tell the truth about tradeoffs, and leave the system better than they found it.

Get Site Reliability Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Sequoia Connect logo

Lead Test Automation Engineer

Mexico
✓ From careers page· 17m ago
Sequoia Connect logo

Java Backend Developer

Mexico City, MX
✓ From careers page· 17m ago
Accenture logo

AI Engineering Consultant

$54k–$206kBoston, MA
✓ From careers page· 29m ago
Accenture logo

Backend Java & Spring Boot Developer

Recife, PE
✓ From careers page· 34m ago

Frequently asked questions

What skills are required for Site Reliability Engineer at Retool?

The required skills for Site Reliability Engineer at Retool include: Kubernetes, Terraform, AWS, PostgreSQL, Networking, Helm, Go, Python, TypeScript, Java, Ruby.

What is the seniority level for Site Reliability Engineer at Retool?

Site Reliability Engineer at Retool is a Mid Level level position.

How do I apply for Site Reliability Engineer at Retool?

You can view the full description and apply for Site Reliability Engineer at Retool on EchoJobs: https://echojobs.io/job/retool-site-reliability-engineer-sre-ukzv7.