Software Engineer, Inference
Location: SF Bay Area, CA, London, UK
Department: Systems Research & Engineering
Location Type: HYBRID
Employment Type: FULL_TIME
About Luma AI
Luma’s mission is to build multimodal AI to expand human imagination and capabilities.
Role & Responsibilities
- Ship new model architectures by integrating them into our inference engine
- Collaborate closely across research, engineering and infrastructure to streamline and optimize model efficiency and deployments
- Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows
- Automate, test and maintain our inference services to ensure maximum uptime and reliability
- Optimize deployment workflows to scale across thousands of machines
- Manage and optimize our inference workloads across different clusters & hardware providers
- Build sophisticated scheduling systems to optimally leverage our expensive GPU resources while meeting internal SLOs
- Build and maintain CI/CD pipelines for processing/optimizing model checkpoints, platform components, and SDKs for internal teams to integrate into our products/internal tooling
Background
- Strong Python and system architecture skills
- Experience with model deployment using PyTorch, Huggingface, vLLM, SGLang, tensorRT-LLM, or similar
- Experience with queues, scheduling, traffic-control, fleet management at scale
- Experience with Linux, Docker, and Kubernetes
- Bonus points:
- Experience with modern networking stacks, including RDMA (RoCE, Infiniband, NVLink)
- Experience with high performance large scale ML systems (>100 GPUs)
- Experience with FFmpeg and multimedia processing
Example Projects
- Create a resilient artifact store that manages all checkpoints across multiple versions of multiple models
- Enable hotswapping of models for our GPU workers based on live traffic patterns
- Build a robust queueing system for our jobs that take into account cluster availability and user priority
- Architect a e2e model serving deployment pipeline for a custom vendor
- Integrate our inference stack into an online reinforcement learning pipeline
- Regression & precision testing across different hardware platforms
- Building a full tracing system to trace the end-to-end lifetime of any inference workload
Tech stack
Must have
- Python
- Redis
- S3-compatible Storage
- Model serving (one of: PyTorch, vLLM, SGLang, Huggingface)
- Understanding of large-scale orchestration, deployment, scheduling (via Kubernetes or similar)
Nice to have
- CUDA
- FFmpeg
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say
