Rehire logo

Junior Site Reliability Engineer

Rehire

On-site
Guadalajara, MX
Full-time
Entry
1+ yrs
Salary not listedPosted 44m ago

Real job — pulled straight from Rehire’s careers page · Verified October 8, 2026 · No reposts.

Job description

Rehire is hiring a Junior Site Reliability Engineer — a full-time, based in Guadalajara, MX role. Apply directly on Rehire's careers page below.

Junior SRE Engineer Guadalajara Mexico

Department: IT

Location: Guadalajara

Employment Type: FullTime

Role Overview

At Rehire, we are partnering with a US-based data engineering and cloud technologies company to find a Junior SRE Engineer to join its Reliability Engineering team. You will be embedded with the SRE function of a financial services client, supporting a regulated consumer-lending platform on AWS with hundreds of microservices and event-driven pipelines, where reliability directly impacts customer trust and compliance.

This is a unique opportunity to grow in an environment where AI is already part of day-to-day operations, including an AI SRE co-pilot, purpose-built AI agents for incident triage and monitoring, and a formal AI governance program. Working under the guidance of senior engineers and architects, you will help operate, improve, and learn from this AI-driven reliability program.

Key Responsibilities:

· Support incident response as a shadow or secondary responder, using AI-driven detection, correlation, and root-cause analysis tools.

· Help build incident timelines from metrics, logs, traces, and deploy history, and document postmortems and corrective actions through to closure.

· Help operate and monitor AI SRE sub-agents (incident summarization, monitor-gap detection, usage attribution), flagging anomalies for senior review.

· Build and maintain Datadog monitors, dashboards, and SLO definitions as code using Terraform, and support SLI/SLO and error-budget reviews for critical customer journeys.

· Help reduce alert noise with AI/ML-assisted detection (anomaly, outlier, and forecast monitors, dynamic thresholds) and run recurring monitor-hygiene reviews.

· Support the day-to-day reliability of AWS workloads (ECS/Fargate, EKS, Lambda, RDS/Aurora, ALB, SQS/SNS, Step Functions).

· Identify capacity, saturation, and cloud cost anomalies, and help attribute spend and telemetry volume to owning teams and services.

· Write automation in Python and Bash against platform APIs (Datadog, AWS, GitHub, PagerDuty, Jira) and contribute Terraform modules through pull requests.

· Help integrate reliability controls and AI-assisted checks into CI/CD pipelines, and create runbooks progressively automated toward self-healing.

· Participate in architecture, reliability, and AI-risk reviews, learning how compliance frameworks (PCI-DSS, SOC 2, SOX, GLBA) apply in a regulated environment.

Requirements:

· Bachelor's degree in Computer Science, Engineering, or a related field.

· 1–3 years of experience in SRE, DevOps, cloud infrastructure, platform, or production-support engineering.

· Advanced English (oral and written). REQUIRED

· Hands-on exposure to AWS (or an equivalent hyperscaler) across compute, networking, storage, and managed database services.

· Exposure to at least one observability platform (Datadog preferred; Grafana/Prometheus, New Relic, CloudWatch, or ELK/OpenSearch also relevant).

· Foundational understanding of SLI, SLO, and error-budget concepts.

· Foundational knowledge of containers and orchestration (Docker, ECS, or Kubernetes) and serverless execution models.

· Beginner-to-intermediate experience with Infrastructure as Code (Terraform preferred; Ansible or CloudFormation acceptable).

· Scripting experience in Python, Bash, or similar, including consuming REST APIs and parsing JSON.

· Basic Linux troubleshooting and networking fundamentals (DNS, TLS, load balancing, timeouts, and retries).

· Comfort with Git, pull-request workflows, and CI/CD tools (GitHub Actions, Jenkins, GitLab CI, ArgoCD, or similar).

· Familiarity with incident management and on-call concepts (severity models, escalation policies, PagerDuty or Opsgenie).

· Experience in product engineering services, enterprise software, or fintech is a plus.

· Awareness of compliance frameworks (PCI-DSS, SOC 2, SOX, GLBA) is a plus.

Key Competencies:

· Automation mindset: you would rather automate a task the second time you do it than the tenth.

· Good judgment to escalate early instead of sitting on an uncertain production signal.

· Clear written communication: you can explain an incident, a metric, or a trade-off to someone who was not in the room.

· Curiosity about LLM-based assistants and agents applied to operations, and about how to verify that their output is correct.

Preferred Certifications (not required):

· AWS Certified Cloud Practitioner or an Associate-level AWS certification.

· HashiCorp Certified: Terraform Associate.

· Datadog Fundamentals or an equivalent observability certification.

· Certified Kubernetes Administrator (CKA) or KCNA.

About the Position:

· Work Schedule: US shift presential at Guadalajara, Mexico.

· Work Modality: Full-time contractor basis.

· Competitive Salary Paid in USD.

· Work Environment: Dynamic and collaborative.

· Professional Growth: Hands-on learning in AI-driven SRE practices and opportunities for career advancement.

If you meet the requirements and are interested in this exciting opportunity, apply at www.rehire.ar/jobs and send us your CV!

Get Site Reliability Engineer jobs like this→

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
LucidLink logo

Software Engineer, Distributed Systems

Sofia
✓ From careers page· 23h ago
Electrolux Group logo

Software Engineer In Test

Bangalore
✓ From careers page· 23h ago
Electrolux Group logo

Senior QA Engineer

Bangalore
✓ From careers page· 23h ago
Infotree Global Solutions logo

Cybersecurity Design Reviewer Architect

Warsaw, Poland
✓ From careers page· 23h ago

Frequently asked questions

What skills are required for Junior Site Reliability Engineer at Rehire?

The required skills for Junior Site Reliability Engineer at Rehire include: AWS, Datadog, Terraform, Python, Bash, Docker, Kubernetes, Git, CI/CD, Linux, Networking, JIRA, ECS, EKS, Lambda, RDS, SQS, PCI DSS, SOC 2.

What is the seniority level for Junior Site Reliability Engineer at Rehire?

Junior Site Reliability Engineer at Rehire is a Entry level position.

How do I apply for Junior Site Reliability Engineer at Rehire?

You can view the full description and apply for Junior Site Reliability Engineer at Rehire on EchoJobs: https://echojobs.io/job/rehire-junior-sre-engineer-guadalajara-mexico-vlq1g.