Braze

Staff Site Reliability Engineer

Toronto, Ontario
Elasticsearch Redis Kafka Kubernetes Docker Shell Go Ruby MongoDB Chef Terraform Python
This job is closed! Check out or
Description

At Braze, we have found our people. We’re a genuinely approachable, exceptionally kind, and intensely passionate crew.

We seek to ignite that passion by setting high standards, championing teamwork, and creating work-life harmony as we collectively navigate rapid growth on a global scale while striving for greater equity and opportunity – inside and outside our organization.

To flourish here, you must be prepared to set a high bar for yourself and those around you. There is always a way to contribute: Acting with autonomy, having accountability and being open to new perspectives are essential to our continued success. Our deep curiosity to learn and our eagerness to share diverse passions with others gives us balance and injects a one-of-a-kind vibrancy into our culture.

If you are driven to solve exhilarating challenges and have a bias toward action in the face of change, you will be empowered to make a real impact here, with a sharp and passionate team at your back. If Braze sounds like a place where you can thrive, we can’t wait to meet you.

Site Reliability Engineers (SREs) are responsible for keeping all internal-facing services and platforms running smoothly. In a nutshell, SREs ensure site uptime. SREs blend sensible system administrators and software engineers who apply sound engineering principles, operational discipline, and mature automation to the environments and infrastructure services we provide. We specialize in systems–whether it be networking, the Linux kernel, or some more specific interest in scaling–algorithms or distributed systems.

Our team helps to improve automation, infrastructure reliability, and empowers Braze’s other engineering teams to leverage the infrastructure products and platforms we create easily. Braze operates at a massive scale with over 3.3 billion monthly active users across our customers, collecting hundreds of billions of data points each month, and sending billions of messages to end-users daily. We use a diverse technology stack rooted in Ruby on Rails, MongoDB, Redis, Kafka, Kubernetes, and more.  As a Site Reliability Engineer at Braze, you will collaborate with your team and consumer engineering teams to continuously improve the infrastructure, automation, and tooling that build internal products from these technologies.

WHAT YOU'LL DO

  • Partner with Braze’s engineering teams on:
    • Architecting products to effectively utilize infrastructure platforms in a scalable, reliable manner
    • Debugging reliability and scalability issues across all stack layers, including the products built using our infrastructure platforms
    • Make monitoring and alerting alerts on symptoms and not on outages
    • Ensure that Braze meets our strict enterprise-grade SLAs with customers
  • Develop Braze’s internal platform infrastructure:
    • Create Infrastructure as code using  Chef, Terraform, and Kubernetes
    • Develop deployment pipelines for applications in multiple languages using Docker, Kubernetes, etc
    • Provide centralized/common tooling, services, and automation frameworks that are critical for scaling operations, capacity management, reducing operational pain, and improving the day-to-day workflow of Braze’s engineering teams
  • Manage incidents:
    • Be on a PagerDuty rotation to respond to availability incidents and provide support for other engineers
    • Use your on-call shift to prevent incidents from ever happening
    • Retrospect everything that happens to turn lessons into system improvements/changes, automation, etc

WHO YOU ARE

  • 10+ years of experience as a Software, DevOps, or Site Reliability Engineer
  • You think about systems - interfaces, boundaries, edge cases, failure modes, behaviors, specific implementations
  • Have an urge to collaborate, document, and deliver quickly
    • Collaborating across the global remote teams, often working asynchronously
    • Document everything so you don't need to learn the same thing (or plan the same work) twice
    • Delivering fast to delight our customers–even internal ones
  • Have an enthusiastic, go-for-it attitude. When you see something broken, you can't help but fix it
  • Have a desire to solve everyday challenges facing software engineers and automate their toil away
  • Have an excellent ability to manage multiple tasks and expectations at once
  • Know your way around Linux and Unix Shell
  • Have strong programming skills
    • We are using Ruby, Python and Go
  • Have experience with Docker, Kubernetes, Terraform, or similar IaC technologies
  • Deep experience in at least one of: Redis, Memcache, Elasticsearch, Prometheus, Datadog

WHAT WE OFFER

From comprehensive benefits to remote availability to flexible time off, we’ve got you covered so you can prioritize work-life harmony.

  • Competitive compensation that may include equity
  • Retirement and Employee Stock Purchase Plans
  • Flexible paid time off
  • Comprehensive benefit plans covering medical, dental, vision, life, and disability
  • Family services that include fertility benefits and equal paid parental leave
  • Professional development supported by formal career pathing, learning platforms, and tuition reimbursement 
  • Community engagement opportunities throughout the year, including an annual company wide Volunteerism Week 
  • Employee Resource Groups that provide supportive communities within Braze
  • Collaborative, transparent, and fun culture recognized as a Great Place to Work® 

Details of these benefit plans will be provided if a candidate receives an offer of employment. Benefits may vary by location.

ABOUT BRAZE

Braze (Nasdaq: BRZE) is a leading comprehensive customer engagement platform that powers interactions between consumers and brands they love. With Braze, global brands can ingest and process customer data in real time, orchestrate and optimize contextually relevant, cross-channel marketing campaigns and continuously evolve their customer engagement strategies.

Braze is proudly certified as a Great Place to Work® in the U.S., the UK and Singapore. We ranked #1 on Great Place to Work UK’s 2023 Best Workplaces (Medium), #3 on Great Place to Work UK’s 2023 Best Workplaces for Wellbeing (Medium), #4 on Great Place to Work’s 2023 Best Workplaces in Europe (Medium), #5 on Fortune’s 2022 Best Workplaces for Millennials in the US, #10 on Great Place to Work UK’s 2023 Best Workplaces for Women (Large), #19 on Fortune’s 2023 Best Workplaces in New York (Large), and were named as a Top Achiever on Great Place to Work UK’s 2023 Best Workplaces in Tech.

You’ll find many of us at headquarters in New York City or around the world in Austin, Berlin, Chicago, Jakarta, London, Paris, San Francisco, Singapore, Sydney and Tokyo – not to mention our employees in nearly 50 remote locations.

Please see our Candidate Privacy Policy for more information on how Braze processes your personal information during the recruitment process and, if applicable based on your location, how you can exercise any privacy rights.

There are more than 50,000 engineering jobs:

Subscribe to membership and unlock all jobs

Engineering Jobs

50,000+ jobs from 4,500+ well-funded companies

Updated Daily

New jobs are added every day as companies post them

Refined Search

Use filters like skill, location, etc to narrow results

Become a member

🥳🥳🥳 250 happy customers and counting...

Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.

Cancel anytime / Money-back guarantee

Wall of love from fellow engineers