Kody logo

Senior Site Reliability Engineer

Kody

On-site
Shenzhen, Guangdong Province, China
Full-time
Senior
Staff
8+ yrs
Salary not listedPosted 1h ago

Real job — pulled straight from Kody’s careers page · Verified September 11, 2026 · No reposts.

Job description

Kody is hiring a Senior Site Reliability Engineer — a full-time, based in Shenzhen, Guangdong Province, China role. Apply directly on Kody's careers page below.

Senior Site Reliability Engineer

Location: Shenzhen, Guangdong Province, China

Department: Technology

Workplace: on_site

Description

Job Summary

Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Shenzhen, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.

Key Responsibilities

  • Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
  • Production Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
  • SLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
  • Continuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
  • Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
  • Technical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.

Requirements

Qualifications & Requirements

  • Experience: 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
  • Core Technical Stack: Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
  • Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
  • Domain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
  • SRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
  • Location & Communication: Based in Hong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.

Leadership & Operational Excellence

  • Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
  • Structured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
  • Crisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
  • Engineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.

Benefits

- Competitive package

- A dynamic and innovative team

- Collaborative, inclusive working environment where your contributions are recognized

Get Site Reliability Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
NERA Economic Consulting logo

Solutions Architect

Cluj-Napoca
✓ From careers page· 15m ago
NERA Economic Consulting logo

Senior Platform Engineer

Cluj-Napoca
✓ From careers page· 16m ago
NERA Economic Consulting logo

IT Systems Engineer

Cluj-Napoca
✓ From careers page· 16m ago
Granica logo

Software Engineer, Infrastructure

Bengaluru
✓ From careers page· 20m ago

Frequently asked questions

What skills are required for Senior Site Reliability Engineer at Kody?

The required skills for Senior Site Reliability Engineer at Kody include: AWS, Kubernetes, Terraform, PostgreSQL, Redis, Kafka, Linux, Datadog, Prometheus, Grafana, PCI DSS, SRE, DevOps, Microservices.

What is the seniority level for Senior Site Reliability Engineer at Kody?

Senior Site Reliability Engineer at Kody is a Senior / Staff level position.

How do I apply for Senior Site Reliability Engineer at Kody?

You can view the full description and apply for Senior Site Reliability Engineer at Kody on EchoJobs: https://echojobs.io/job/kody-senior-site-reliability-engineer-9po2i.