Lead Infrastructure and Reliability Engineer (Systems & Scale)
Location: SF Bay Area, CA
Department: Product Engineering
Location Type: HYBRID
Employment Type: FULL_TIME
About Luma AI
Where You Come In
- Kernels
- Containers
- Schedulers
- Networking
- Storage
- GPU behavior
What You’ll Own
Reliability of the Frontier
- Architect and operate large, heterogeneous GPU environments under extreme demand
- Improve utilization and performance where small gains materially change company outcomes
- Resolve failures that span hardware, OS, runtimes, and orchestration
- Eliminate entire classes of instability
- Build mechanisms that make heroics unnecessary
Scaling Training & Inference
- Define how infrastructure and workloads evolve as cluster size and concurrency grow
- Design scheduling, placement, and resource management approaches for increasingly complex jobs
- Work directly with research to build the systems required for new model capabilities
- Ensure inference platforms scale rapidly without sacrificing reliability or latency
- Anticipate where today’s abstractions will fail and redesign ahead of them
Building the Organization
- Hire and develop exceptional systems and reliability engineers
- Set the bar for technical depth, judgment, and production ownership
- Shape architecture early through strong partnerships with research and product
- Translate reliability constraints into long-term platform strategy
Who You Are
Required:
- Deep expertise in Linux and distributed systems
- Experience operating GPU / accelerator clusters in real production environments
- Strong fluency in Kubernetes and modern open-source infrastructure
- Comfortable debugging across hardware → kernel → runtime → orchestration
- You understand how systems behave under contention and at scale
- You write code and build automation
- You think in bottlenecks, failure modes, and tradeoffs
- Engineers trust your judgment, especially when things break
Leadership Expectations
- You raise reliability standards across the company
- You influence product and research architecture early
- You build strong partnerships, not ticket queues
- You attract and level up exceptional engineers
- You are curious how models use infrastructure, because improving systems expands what becomes possible
Why This Role Is Special
- How research progresses
- How products scale
- How customers trust us
- And how the engineering organization grows
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say
