
Real job — pulled straight from Apply Digital’s careers page · Verified September 6, 2026 · No reposts.
Job description
Apply Digital is hiring a Site Reliability Engineer (Remote) — a full-time, remote role. Apply directly on Apply Digital's careers page below.
Site Reliability Engineer (SRE)
Team: Managed Services Team
Location: Santiago, Latin America
Commitment: Full-Time Permanent
Workplace Type: remote
Apply Digital is looking for an SRE Engineer to join our globally distributed team. This is a hybrid SRE/Service Desk role designed for someone who is passionate about reliability engineering and comfortable supporting day-to-day operational needs across multiple UK e-commerce clients.
You will be a key contributor in maintaining the health and performance of client platforms, responding to incidents, and continuously improving observability and operational processes. While your primary focus is SRE, you will collaborate closely with the Service Desk team to support triaging, escalation, and resolution workflows.
This role is ideal for someone with 2–3 years of experience who thrives in a fast-paced, multi-client environment, values clear documentation, and is comfortable working with a high degree of autonomy during their shift.
WHAT YOU’LL DO
-
Monitor platform health across multiple client environments using tools like Grafana and Prometheus, or other monitoring tools
-
Respond to and triage incidents, following established runbooks and escalation paths
-
Participate in post-incident reviews and contribute to postmortem documentation
-
Support the Service Desk team with technical triaging, incident classification, and resolution
-
Maintain and improve observability dashboards, alerts, and SLI, and SLO tracking
-
Write and maintain runbooks, operational documentation, and knowledge base articles
-
Identify recurring issues and propose automation or process improvements to reduce toil
-
Participate in on-call rotation covering weekends (alternating schedule — one weekend on, one weekend off)
-
Collaborate with the EMEA team during shift overlap to ensure smooth handoffs and continuity
-
Support root cause analysis and contribute to continuous improvement initiatives
WHAT WE’RE LOOKING FOR:
-
Strong proficiency in English (written and verbal communication) is required
-
2–3 years of experience in SRE, platform operations, or a technical Service Desk role
-
Experience with monitoring and observability tools such as Grafana, Prometheus, or equivalent
-
Solid understanding of incident management processes (triaging, escalation, postmortems)
-
Experience supporting e-commerce platforms
-
Scripting skills in shell and/or Python for automation and operational tasks
-
Familiarity with containerization concepts (Docker, Kubernetes) at an operational level
-
Experience working in Agile environments and using ticketing tools (e.g. Jira)
-
Comfort working independently during early-morning shifts with minimal supervision
-
Strong documentation habits and attention to detail
-
Experience with Agile processes, testing, and code review
-
Strong experience with scripting - shell, Python, etc.
-
Excellent customer service attitude, communication skills (written and verbal), and interpersonal skills
-
Excellent analytical and problem-solving skills
-
Ability to communicate effectively with technical and non-technical stakeholders. You should feel comfortable explaining technical concepts in simple terms
-
Experience working in fast-paced, Agile environments, balancing priorities across multiple projects
-
Experience with Google Cloud Platform (GCP) or other major cloud providers (AWS, Azure)
-
Familiarity with CI/CD pipelines (GitHub Actions, GitLab CI)
-
Basic experience with Infrastructure as Code tools such as Terraform
-
Basic knowledge of AIOps concepts and their application in operational workflows
-
SRE or cloud certifications (Google Cloud, AWS, Kubernetes)
NICE TO HAVE:
Get Site Reliability Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs

Staff Site Reliability Engineer, Government (Remote)


Site Reliability Engineer, Database Administration (Remote)

Frequently asked questions
Is Site Reliability Engineer (Remote) at Apply Digital a remote job?
Yes, Site Reliability Engineer (Remote) at Apply Digital is a remote position. Candidates in Santiago, CL may be preferred.
What skills are required for Site Reliability Engineer (Remote) at Apply Digital?
The required skills for Site Reliability Engineer (Remote) at Apply Digital include: Grafana, Prometheus, Shell, Python, Docker, Kubernetes, JIRA, Agile, GCP, AWS, Azure, GitHub Actions, GitLab CI, Terraform.
What is the seniority level for Site Reliability Engineer (Remote) at Apply Digital?
Site Reliability Engineer (Remote) at Apply Digital is a Mid Level level position.
How do I apply for Site Reliability Engineer (Remote) at Apply Digital?
You can view the full description and apply for Site Reliability Engineer (Remote) at Apply Digital on EchoJobs: https://echojobs.io/job/apply-digital-site-reliability-engineer-sre-dtbld.