
Real job — pulled straight from SkanAI’s careers page · Verified August 7, 2026 · No reposts.
Job description
SkanAI is hiring a Site Reliability Engineer — a full-time, based in Bengaluru, India role. Apply directly on SkanAI's careers page below.
Sr Site Reliability Engineer
Location: Bengaluru, India
Department: AI Solutions Engineering & Support
Experience: 5-8
- Monitor platform health and performance across all cloud-hosted customer environments using observability tooling (Prometheus, Grafana, Datadog, or equivalent)
- Respond to and own P1/P2 incidents — lead triage, diagnosis, and resolution; drive MTTR reduction through structured post-incident review
- Perform ongoing capacity and environment planning for customer cloud deployments — anticipating growth and preventing resource-related incidents
- Design and implement SRE automation to eliminate repetitive operational toil — agentic alerting, auto-remediation scripts, and automated runbook execution
- Manage change and release events that affect production customer environments — coordinating with DevOps and product teams to minimize risk
- Maintain and improve runbooks for all known failure patterns and operational procedures
- Contribute to HA/DR playbook validation
- Participate in on-call rotation and respond to alerts within defined SLA windows
- Track and report on SLO/SLA adherence — contributing to monthly operational reports and QBR data
- Identify and escalate environment risks proactively — before they become customer-facing incidents
- Collaborate with the Automation Engineering team to develop agentic workflows that automate triage, routing, and remediation
- Incident response records and post-mortems for all P1/P2 events — with root cause, remediation steps, and prevention actions documented
- SLA/SLO dashboards: availability and performance reports for all cloud customer environments — updated continuously
- Capacity and environment sizing plans per customer — reviewed quarterly or on significant usage change
- Automated toil-reduction scripts and agentic remediation workflows — measurable reduction in manual operational hours
- Runbooks for all known failure patterns — maintained and validated against real incidents
- Change management records for all production environment events
- Target: ≥ 99.9% uptime across all cloud-hosted customer environments
- 5-8 years of experience as a SRE engineer
- Cloud platforms: AWS, Azure, or GCP — environment management, networking, IAM, and observability at scale
- Observability and monitoring: Prometheus, Grafana, Datadog, or equivalent — building dashboards, alerts, and SLO tracking
- Infrastructure as Code: Terraform, Ansible, or Pulumi — provisioning and configuration management
- Incident management: on-call discipline, structured MTTR mindset, post-mortem culture, and blameless review practices
- Scripting and automation: Python, Bash — automation of operational tasks and agentic workflow development
- Linux systems administration: process management, log analysis, performance tuning
- Container orchestration: Kubernetes and Docker — deployment management and debugging in production
- SRE fundamentals: SLI/SLO/SLA definition, error budget management, toil measurement and reduction
- Strong written documentation skills — clear, evidence-based runbooks and incident reports
Get Site Reliability Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs
Frequently asked questions
What skills are required for Site Reliability Engineer at SkanAI?
The required skills for Site Reliability Engineer at SkanAI include: AWS, Azure, GCP, Prometheus, Grafana, Datadog, Terraform, Ansible, Python, Bash, Linux, Kubernetes, Docker, IAM, CI/CD.
What is the seniority level for Site Reliability Engineer at SkanAI?
Site Reliability Engineer at SkanAI is a Senior level position.
How do I apply for Site Reliability Engineer at SkanAI?
You can view the full description and apply for Site Reliability Engineer at SkanAI on EchoJobs: https://echojobs.io/job/skan-ai-sr-site-reliability-engineer-4zouj.

