Principal Site Reliability Engineer
Team: Cloud Services Engineering
Location: Bengaluru
Commitment: Full-Time
Workplace Type: hybrid
We’re building an AI-native platform powered by LLMs and autonomous agents, designed to scale intelligently and operate with minimal human intervention. Our stack runs on AWS + Kubernetes, with deep integration of LLM capabilities to create self-healing, adaptive systems.
As a Principal SRE Engineer, you’ll define the reliability vision and architecture of our platform. This is a hands-on leadership role where you’ll drive strategy, design large-scale distributed systems, and pioneer AI-driven operations (AIOps).
You won’t just improve reliability—you’ll redefine how reliability works in an AI-first world.
WHAT YOU BRING
- 10+ years in SRE / Platform / Distributed Systems Engineering.
- Proven experience designing and operating large-scale distributed systems.
- Deep expertise in:
- AWS architecture at scale
- Kubernetes internals and operations
- System reliability, scalability, and performance engineering
- Strong programming skills (Python / Go) with focus on building platforms/tools.
- Experience leading cross-functional technical initiatives.
- Ability to influence architecture and engineering culture across the company.
- Experience integrating LLMs into production systems (e.g., via OpenAI API).
- Built or designed AI-driven automation / AIOps systems.
- Familiarity with:
- Agent frameworks (LangChain, AutoGen)
- RAG pipelines and vector databases
- Strong interest in autonomous systems and self-healing infrastructure.
AI / LLM & Future Systems (Highly Valued)
WHAT YOU WILL BE DOING
- Own and define the long-term reliability strategy and architecture.
- Design planet-scale, highly resilient systems on AWS and Kubernetes (EKS).
- Lead the development of autonomous operations platforms powered by AI agents.
- Architect and implement LLM-driven SRE systems using tools like the OpenAI API:
- Intelligent incident detection and triage
- Automated root cause analysis
- Self-healing remediation systems
- Establish gold standards for SRE practices:
- SLOs, SLAs, error budgets
- Incident management frameworks
- Reliability-first system design
- Drive observability architecture at scale (metrics, logs, traces, events).
- Lead cross-team initiatives to embed reliability into product and platform design.
- Mentor senior engineers and act as a technical authority across teams.
- Guide decisions around cost optimization, scalability, and performance trade-offs.
- Introduce chaos engineering and resilience testing at system-wide level.
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say
