
Real job — pulled straight from AQEMIA’s careers page · Verified September 27, 2026 · No reposts.
Job description
AQEMIA is hiring a Site Reliability Engineer — a full-time, based in Paris role. Apply directly on AQEMIA's careers page below.
SRE
Team: Core, Security & Infra
Location: Paris
Workplace Type: hybrid
About AQEMIA
About our Team
About our Engineering Department
The role
As an SRE within our Core Engineering team, you’ll help operate and evolve Aqemia’s cloud platform end to end — from infrastructure-as-code and GitOps delivery to observability, security and FinOps.
The Core team owns the platform, and you’ll work to make it reliable, scalable, secure and easy to use. At Aqemia, nothing is deployed by hand and infrastructure changes flow through Git, so automation and reproducibility are central to how the team works.
The scale at Aqemia is different from a typical product company: bursty, large-scale parallel scientific computation, GPU fleets to plan and optimize, and some of the company's most valuable data to protect, all while keeping cost under control.
You’ll have the opportunity to contribute directly to architectural and tooling decisions, take ownership of meaningful parts of the platform, and work closely with the teams that depend on it every day. In a small Engineering organisation, improvements to the platform have a direct impact on how quickly our scientists and engineers can work.
As the platform evolves, from today's GitOps pipelines toward productionized MLOps and autonomous discovery workflows, you'll grow with it, taking on more scope rather than staying in a fixed lane.
Responsibilities
- Own AWS infrastructure end to end, built and maintained as code with OpenTofu and Terragrunt - no manual changes, no exceptions.
- Operate and evolve Kubernetes workloads via GitOps (ArgoCD, Helm, Kustomize), and drive adoption of standardized infrastructure patterns across teams.
- Contribute to the platform's reliability practice: own observability (metrics, logs, alerting), respond to incidents, and run blameless postmortems through to completed action items.
- Manage cloud cost as a shared responsibility - producing the monthly cost report, maintaining the cost allocation model, and partnering with teams to plan capacity ahead of large GPU compute campaigns.
- Set and enforce security posture across the platform: patching, vulnerability follow-up, and least-privileged access management.
- Build internal tooling and CI/CD pipelines that reduce friction for engineering, ML and scientific teams, and make the platform approachable to non-infrastructure users.
- Shape platform architecture and long-term strategy through technical reviews and sprint planning, sharing knowledge across infrastructure and DevOps topics.
Qualifications
- Strong platform/infrastructure engineering background, with 2-3+ years of experience post-degree.
- Deep hands-on expertise in AWS and infrastructure-as-code (Terraform/OpenTofu, Terragrunt).
- Strong production experience with Kubernetes and GitOps delivery (ArgoCD, Helm, Kustomize).
- Experience building and maintaining CI/CD pipelines (GitHub Actions or GitLab).
- Experience with cloud security practices and least-privileged access management.
Nice-to-have
- MLOps experience - training/inference pipelines, model lifecycle, workflow orchestrators.
- GPU capacity planning - autoscaling GPU fleets, spot strategies, quota management.
- Exposure to AI-driven or data-intensive workflows.
- Experience with another cloud provider beyond AWS (e.g. GCP).
Our recruitment process
- First discussion with our Talent Acquisition
- Hiring Manager’s interview: you’ll meet directly with your future manager
- Technical assessment of your skills in a deep-dive interview with the team
- Cultural fit interview with our co-founder and COO, Emmanuelle
- Final interview with our co-founder and CEO, Maximillien
Why Join Us?
Get Site Reliability Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Site Reliability Engineer at AQEMIA?
The required skills for Site Reliability Engineer at AQEMIA include: AWS, Kubernetes, ArgoCD, Helm, CI/CD, GitHub Actions, GitLab, MLOps, Terraform.
What is the seniority level for Site Reliability Engineer at AQEMIA?
Site Reliability Engineer at AQEMIA is a Mid Level level position.
How do I apply for Site Reliability Engineer at AQEMIA?
You can view the full description and apply for Site Reliability Engineer at AQEMIA on EchoJobs: https://echojobs.io/job/aqemia-sre-fkwbm.