
Real job — pulled straight from RapidAI’s careers page · Verified September 9, 2026 · No reposts.
Job description
RapidAI is hiring a Senior Site Reliability Engineer — a full-time, based in Bangalore, India role. Apply directly on RapidAI's careers page below.
Senior Site Reliability Engineer
Team: Engineering
Location: Bangalore, India
Commitment: Full Time
Workplace Type: hybrid
- Own the availability, performance, and incident response for Rapid's production EKS clusters
- Design and operate the full observability stack — metrics, logs, traces — with
Open Telemetry as the foundation - Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
- Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
- Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
- Tune autoscaling, networking, and cost efficiency across AWS workloads
- On-call rotation with the expectation you'll also fix the underlying cause, not just the alert
- 10+ years in SRE, DevOps, or infrastructure engineering roles
- Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and the
surrounding ecosystem - Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic
- Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)
- Strong foundation in Linux, networking, and distributed systems fundamentals
- Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents)
Comfortable writing automation in Go, Python, or Bash — you reach for code when the GUI runs out - Startup mindset: you make decisions with incomplete information and iterate quickly
What You Do:
Get Site Reliability Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs



Senior SDET - Test Automation & Performance Engineering (Remote)

Frequently asked questions
What skills are required for Senior Site Reliability Engineer at RapidAI?
The required skills for Senior Site Reliability Engineer at RapidAI include: AWS, EKS, EC2, IAM, RDS, S3, CloudWatch, Kubernetes, OpenTelemetry, Terraform, Helm, Linux, Networking, Go, Python, Bash, Prometheus, Grafana.
What is the seniority level for Senior Site Reliability Engineer at RapidAI?
Senior Site Reliability Engineer at RapidAI is a Senior level position.
How do I apply for Senior Site Reliability Engineer at RapidAI?
You can view the full description and apply for Senior Site Reliability Engineer at RapidAI on EchoJobs: https://echojobs.io/job/rapidai-senior-site-reliability-engineer-4m945.