
Real job — pulled straight from Index Exchange’s careers page · Verified September 10, 2026 · No reposts.
Job description
Index Exchange is hiring a Staff Platform Site Reliability Engineer — a full-time, based in Toronto, ON role. Apply directly on Index Exchange's careers page below.
Staff Platform Site Reliability Engineer
Location: Toronto (8 Spadina Ave)
Department: Technology Operations
Location Type: HYBRID
Employment Type: FULL_TIME
The Role
What You'll Work On
- Build the platform, not run it. You'll design and deliver multi-tenant Kubernetes infrastructure across bare-metal and public cloud, build infrastructure-as-code frameworks that push changes across thousands of servers, and write the standard libraries, SDKs, and platform APIs that every engineering team at Index Exchange depends on. This is the work that takes Index Cloud to the next level.
- Solve hard distributed systems problems. Real-time bidding with sub-millisecond overhead. Multi-datacenter consistency. Deploying safely to a fleet of thousands. Load balancing at a scale where the easy approaches don't work. If the phrase "globally distributed auction system" makes you lean in, keep reading.
- Own the architecture. You'll drive technical direction through RFCs and design reviews, set standards for shared platform domains, and make the calls on tooling, security posture, and system design. This is a Staff-level role—you'll influence engineering direction across multiple divisions.
- Make engineers faster. Build golden paths, self-service tooling, and platform APIs that let engineering teams ship without friction. The best platform work is invisible to its users—it just works.
- Mentor and raise the bar. Coach engineers, foster a culture of engineering excellence, and collaborate across Cloud Platform Operations, SRE, Network, Security, and Software Engineering teams.
What You Bring
Must Have
- 8+ years in platform engineering, SRE, infrastructure engineering, or DevOps.
- Deep experience with Linux internals: kernel tuning, network stack, system observability, security.
- Strong Kubernetes expertise: cluster lifecycle, networking, storage, RBAC, multi-cluster—across bare-metal and cloud (EKS, GKE).
- Infrastructure-as-code at scale: Terraform, Ansible, GitOps (ArgoCD or similar).
- Proficiency in Go, Python, or both—for building libraries, SDKs, and platform APIs, not just scripts.
- Solid networking fundamentals (L2-L7), load balancing, DNS, service discovery.
- A track record of driving technical strategy across teams—not just executing within one.
Valuable Experience
- Distributed storage systems (e.g. Ceph)
- Big data infrastructure: Hadoop, Spark, HBase, Kafka.
- Observability stack design: Prometheus, Grafana, ELK, Mimir, Loki, Tempo.
- Secrets management (Vault), certificate management, access control at scale.
- Hybrid cloud architectures: federating public cloud (AWS, GCP) with on-premises environments.
- Experience with bare-metal infrastructure in globally distributed data centers
Who Thrives Here
- Sees a complex, distributed system and wants to understand how it breaks.
- Treats operational readiness as a first-class engineering concern, not an afterthought.
- Prefers building the right abstraction over fighting the same fire twice.
- Can drive alignment across teams without needing a title to do it.
- Finds it genuinely rewarding to make other engineers more productive.
- Comprehensive health, dental, and vision plans for you and your dependents
- Paid time off, health days, and personal obligation days plus flexible work schedules
- Competitive retirement matching plans
- Equity packages
- Generous parental leave available to birthing, non-birthing, and adoptive parents
- Annual well-being allowance plus fitness discounts and group wellness activities
- Commuter benefits and discounts, where available
- Employee assistance program
- Mental health first aid program that provides an in-the-moment point of contact and reassurance
- One day of volunteer time off per year and a donation-matching program
- Monthly town halls and regular community-led team events
- Multiple resources and programming to support continuous learning
- A workplace that supports a diverse, equitable, and inclusive environment – learn more here
Get Site Reliability Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Staff Platform Site Reliability Engineer at Index Exchange?
The required skills for Staff Platform Site Reliability Engineer at Index Exchange include: Linux, Kubernetes, Terraform, Ansible, Go, Python, Prometheus, Grafana, Elasticsearch, AWS, GCP, Hadoop, Spark, Kafka.
What is the seniority level for Staff Platform Site Reliability Engineer at Index Exchange?
Staff Platform Site Reliability Engineer at Index Exchange is a Staff level position.
How do I apply for Staff Platform Site Reliability Engineer at Index Exchange?
You can view the full description and apply for Staff Platform Site Reliability Engineer at Index Exchange on EchoJobs: https://echojobs.io/job/index-exchange-staff-platform-site-reliability-engineer-ayx3s.