
Senior Site Reliability Engineer, Voice AI Platform
Real job — pulled straight from Skit.ai’s careers page · Verified July 12, 2026 · No reposts.
Job description
Skit.ai is hiring a Senior Site Reliability Engineer, Voice AI Platform — a full-time, based in Bangalore, India role. Apply directly on Skit.ai's careers page below.
Senior Site Reliability Engineer — Voice AI Platform
Location: Bangalore, India
Department: Technology
Experience: 5+ Years
- SLOs and error budgets. Define and defend service-level objectives for availability and latency across the call path, and use error budgets to steer the balance between shipping and stability.
- Observability. Own the metrics, tracing, and logging stack so failures surface fast and root cause is minutes not hours — distributed traces across signaling, ASR, LLM, TTS, and infra, with dashboards and alerting that page on real problems and stay quiet otherwise.
- The real-time media path. Keep SIP signaling and RTP media healthy at scale — concurrency, jitter, packet loss, session setup — and the reliability of the components that carry them.
- Capacity and autoscaling. Plan for peak (campaign windows that push toward the platform's concurrency ceiling), pre-warm capacity ahead of demand, and tune autoscaling so we neither drop calls nor burn money idling.
- Incident response. Run a calm, structured on-call: triage, mitigation, clear comms to stakeholders on regulated accounts, and blameless postmortems that actually change the system.
- Multi-cloud resilience. Design for failure across AWS, GCP, and Azure — redundancy, failover, disaster recovery, and the data-residency constraints that come with Indian banking and telecom clients.
- Automation and toil reduction. Turn manual operations into infrastructure-as-code and self-healing systems. If you did it twice by hand, the third time is a script.
- First 90 days. Learn the call path end to end. Establish baseline SLIs for availability and latency, close the biggest gaps in alerting, and take a full turn in the on-call rotation.
- By 6 months. Published SLOs with error budgets for the core services. A tracing/dashboards setup that makes the ASR→LLM→TTS latency budget visible per call. A repeatable pre-warm-and-scale playbook for campaign peaks.
- By 12 months. Demonstrable reduction in incident frequency and time-to-mitigate. Tested multi-cloud failover for a critical path. On-call toil measurably down through automation.
- Several years running high-availability, high-throughput production systems, including real on-call ownership.
- Depth in at least one major cloud (AWS, GCP, or Azure) and with containers/Kubernetes.
- Strong observability practice — metrics, distributed tracing, and logging (e.g. Prometheus/Grafana, OpenTelemetry, Tempo/Jaeger).
- Fluency with SLIs/SLOs/error budgets and structured incident management.
- Infrastructure-as-code (Terraform or similar) and CI/CD.
- A programming language for automation and tooling (Python, Go, or similar) — beyond shell scripting.
- Solid Linux systems and networking fundamentals; capacity planning and performance tuning.
- Real-time media or VoIP experience — SIP/RTP, media servers, LiveKit, SBC/Kamailio.
- Reliability of GPU/ML serving infrastructure.
- Regulated-industry operations — uptime SLAs, DR, data residency.
- Load testing and chaos engineering at scale.
- PostgreSQL operations at scale.
Get Site Reliability Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs

Manager, Production Operations and Site Reliability Engineering



Frequently asked questions
What skills are required for Senior Site Reliability Engineer, Voice AI Platform at Skit.ai?
The required skills for Senior Site Reliability Engineer, Voice AI Platform at Skit.ai include: SRE, Kubernetes, Prometheus, Grafana, OpenTelemetry, Terraform, Python, Go, Linux, Networking, PostgreSQL, AWS, GCP, Azure, CI/CD.
What is the seniority level for Senior Site Reliability Engineer, Voice AI Platform at Skit.ai?
Senior Site Reliability Engineer, Voice AI Platform at Skit.ai is a Senior level position.
How do I apply for Senior Site Reliability Engineer, Voice AI Platform at Skit.ai?
You can view the full description and apply for Senior Site Reliability Engineer, Voice AI Platform at Skit.ai on EchoJobs: https://echojobs.io/job/skit-ai-senior-site-reliability-engineer-voice-ai-platform-muqkm.