
Real job — pulled straight from Retool’s careers page · Verified August 2, 2026 · No reposts.
Job description
Retool is hiring a Site Reliability Engineer — a full-time, based in San Francisco, CA role. Apply directly on Retool's careers page below.
Site Reliability Engineer (SRE)
Location: San Francisco, New York
Department: Engineering
Location Type: IN_OFFICE
Employment Type: FULL_TIME
- Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations.
- Build the automation that turns today's manual infrastructure work into repeatable systems: Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps.
- Improve observability for Retool Cloud, self-hosted customers, and internal operators. We care less about exposing every metric and more about turning health signals into clear status, likely causes, and recommended actions.
- Design safer deployment, upgrade, and rollback paths so Cloud and managed customers can stay current
- Help move customers from legacy or less-supported deployment models toward supported paths such as Retool's official deployment paths (Blueprints, Kubernetes, and Helm), with migration flows that are repeatable enough for customers, Support, and TAMs to trust.
- Partner with product engineers on infrastructure requirements for new Retool products, especially when they introduce new dependencies
- Lead through ambiguity, make careful risk calls, and communicate clearly while things are moving quickly.
- Write the docs, runbooks, design notes, and migration guides that make complex systems understandable to other engineers and to customers.
- Deep experience operating production infrastructure in AWS.
- Experience improving reliability for customer-facing SaaS systems.
- Strong Kubernetes fundamentals.
- Real Terraform or infrastructure-as-code experience.
- Good operational judgment around databases, especially Postgres.
- Experience building or operating observability systems.
- Programming ability in a language such as Go, Python, TypeScript, Java, or Ruby.
- A bias toward automation. If you find yourself doing the same operational task twice, you should start thinking about the interface, workflow, or tool that eliminates the third time.
- Clear written communication.
- Comfort working directly with customer-facing teams and, when useful, customers themselves.
Get Site Reliability Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Site Reliability Engineer at Retool?
The required skills for Site Reliability Engineer at Retool include: Kubernetes, Terraform, AWS, PostgreSQL, Networking, Helm, Go, Python, TypeScript, Java, Ruby.
What is the seniority level for Site Reliability Engineer at Retool?
Site Reliability Engineer at Retool is a Mid Level level position.
How do I apply for Site Reliability Engineer at Retool?
You can view the full description and apply for Site Reliability Engineer at Retool on EchoJobs: https://echojobs.io/job/retool-site-reliability-engineer-sre-ukzv7.