Staff Site Reliability & DevOps Engineer - Observability
Location: Remote - Hungary; Sofia, Bulgaria
Department: 710 - Software Product Engr
• Design, build, and operate observability platforms based on Grafana and Prometheus
• Define and maintain metrics standards, dashboards, alerts, and SLOs
• Improve signal quality: reduce alert noise, tune thresholds, and improve runbooks
• Support incident response by providing actionable telemetry and post-incident analysis
• Integrate metrics, logs, and traces across distributed systems
• Work with engineering teams to instrument services correctly
• Automate observability configuration using infrastructure as code
• Contribute to reliability improvements through capacity planning and performance analysis
• Required skills and experience
• Strong experience with Prometheus (scraping, federation, recording rules, alerting)
• Strong experience with Grafana (dashboards, alerting, templating, RBAC)
• Solid Linux and networking fundamentals
• Experience running observability stacks in Kubernetes environments
• Infrastructure as code experience (Terraform preferred)
• Familiarity with incident management and on-call practices
• Ability to debug production systems using metrics and logs
Nice to have:
• Experience with logs and traces (e.g. Loki, Tempo, OpenTelemetry)
• Experience operating large-scale or multi-cluster Kubernetes platforms
• Experience with cloud platforms (GCP, AWS, OCI)
• Exposure to SRE concepts such as error budgets and SLO-driven prioritisation
What success looks like
• Engineers trust dashboards and alerts to reflect system health
• Incidents are detected earlier and diagnosed faster
• Alert fatigue is reduced and on-call quality improves
• Observability is treated as a first-class platform capabilit
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say
