Member of Technical Staff — Training
Location: Palo Alto
Department: engineer
About the Role
RadixArk is seeking a Member of Technical Staff — Training to build and scale the systems that train frontier AI models.
You will work on large-scale distributed training infrastructure for LLMs and generative models, pushing the limits of scale, efficiency, and reliability across thousands of GPUs. This role sits at the intersection of ML, systems, and performance engineering.
Your work will directly impact how next-generation AI models are trained and scaled.
This is a deeply technical, high-impact role for engineers who enjoy solving hard systems problems at extreme scale.
Requirements
-
5+ years of experience in ML systems, distributed systems, or large-scale training infrastructure
-
Strong experience with large-scale distributed training (data, tensor, and pipeline parallelism)
-
Deep understanding of GPU/TPU architecture and performance trade-offs
-
Strong knowledge of PyTorch or JAX distributed training stacks
-
Experience debugging performance and stability issues in large training jobs
-
Solid distributed systems fundamentals (networking, consensus, fault tolerance)
-
Proficiency in Python plus a systems language (C++, Go, or Rust)
-
Experience operating production ML systems at scale
Strong Plus
-
Experience training multi-billion-parameter models
-
Familiarity with DeepSpeed, Megatron-LM, FSDP, or custom training stacks
-
Experience with RDMA, InfiniBand, or high-speed interconnects
-
Background in HPC or performance-critical computing
-
Contributions to ML systems open-source projects
-
Experience with checkpointing, fault recovery, and elastic training
-
Experience optimizing training cost efficiency at scale
Responsibilities
-
Design and operate large-scale distributed training systems
-
Optimize throughput, scalability, and hardware efficiency
-
Improve reliability and fault tolerance for long-running training jobs
-
Develop training frameworks and infrastructure tooling
-
Collaborate with model researchers to support frontier experiments
-
Debug and resolve cross-layer performance bottlenecks
-
Build observability systems for training performance and reliability
-
Drive capacity planning and cluster utilization strategies
-
Contribute to long-term training infrastructure architecture
About RadixArk
RadixArk is an infrastructure-first AI company built by engineers who have shipped production AI systems, created SGLang (20K+ GitHub stars, the fastest open LLM serving engine), and developed Miles, our large-scale RL framework.
We build world-class infrastructure for AI training and inference and partner with frontier AI teams and cloud providers.
Our team has coordinated training across 10,000+ GPUs and optimized kernels serving billions of tokens daily.
Join us in building the infrastructure that trains the next generation of AI.
Compensation
We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level.
Equal Opportunity
RadixArk is an Equal Opportunity Employer and welcomes candidates from all backgrounds.
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say
