
Real job — pulled straight from TechBlocks’s careers page · Verified June 18, 2026 · No reposts.
Job description
TechBlocks is hiring a Observability and Chaos Engineering Specialist — a full-time role. Apply directly on TechBlocks's careers page below.
Observability & Chaos Engineering Specialist
Location: Hyderabad, India
Department: Z- S&P SPGI ES Tech
Experience: 7-9
Skills: MCP, observability, chaos engineering
- Design and implement observability frameworks for AI/agent-based systems and distribute cloud-native applications
- Configure and manage Langfuse for LLM/AI workflow observability, tracing, monitoring, and evaluation
- Develop monitoring and telemetry solutions for MCP agent setups and multi-agent orchestration environments
- Implement and optimize AWS native observability services, including:
- Establish centralized logging, distributed tracing, metrics collection, and alerting mechanisms
- Design and execute Chaos Engineering experiments using AWS Fault Injection Simulator (FIS) to validate system resilience and recovery capabilities
- Simulate infrastructure, network, and service failures to identify system weaknesses and improve fault tolerance
- Collaborate with DevOps, Platform Engineering, AI Engineering, and Security teams to improve operational reliability
- Build dashboards, alerts, and health monitoring systems for proactive incident detection and response
- Analyze system behavior under stress conditions and recommend architecture improvements
- Support incident troubleshooting, root cause analysis, and reliability optimization initiatives
- Maintain technical documentation for observability architecture, chaos testing scenarios, and operational runbooks
- 7-9 years of experience in Observability Engineering, SRE, DevOps, or Platform Engineering
- Langfuse for AI/LLM observability
- AI workflow tracing and telemetry
- CloudWatch
- AWS X-Ray
- CloudTrail
- AWS monitoring and logging services
- Hands-on experience with Chaos Engineering practices
- Expertise using AWS Fault Injection Simulator (FIS) for resilience and fault-tolerance testing
- Familiarity with containerized and cloud-native environments (ECS/EKS/Kubernetes)
- Experience with CI/CD pipelines and infrastructure automation
- Strong scripting/programming skills in Python or similar languages
- Strong analytical, troubleshooting, and problem-solving skills
- Experience with:
- OpenTelemetry
- Grafana
- Prometheus
- ELK/OpenSearch stack
- Familiarity with:
- AI/LLM platforms and agentic architectures
- Event-driven and microservices-based systems
- Knowledge of:
- DevSecOps and cloud security monitoring
- Performance engineering and load testing
- AWS certifications preferred
- Experience working in highly regulated or enterprise-scale environments
- Opportunity to work on global projects and Fortune 500 clients
- Exposure to cutting-edge technologies
- Strong learning, mentorship, and career growth programs
- Collaborative and innovation-driven work culture
Get Observability and Chaos Engineering Specialist jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Observability and Chaos Engineering Specialist at TechBlocks?
The required skills for Observability and Chaos Engineering Specialist at TechBlocks include: AWS, CloudWatch, Python, Kubernetes, ECS, EKS, CI/CD, SRE, DevOps, OpenTelemetry, Grafana, Prometheus.
What is the seniority level for Observability and Chaos Engineering Specialist at TechBlocks?
Observability and Chaos Engineering Specialist at TechBlocks is a Senior level position.
How do I apply for Observability and Chaos Engineering Specialist at TechBlocks?
You can view the full description and apply for Observability and Chaos Engineering Specialist at TechBlocks on EchoJobs: https://echojobs.io/job/techblocks-observability-chaos-engineering-specialist-c16jo.