Software Engineer, ML Platform
Location: SF Bay Area, CA, London, UK
Department: Systems Research & Engineering
Location Type: HYBRID
Employment Type: FULL_TIME
- Architect end-to-end model serving pipelines and integrate new model architectures from our research team into our core, high-throughput inference engine.
- Build robust and sophisticated scheduling systems to manage jobs based on cluster availability and user priority, ensuring we optimally leverage thousands of expensive GPU resources.
- Design and implement dynamic, traffic-based systems for hotswapping models on our GPU workers to maximize fleet efficiency and meet product SLOs.
- Own the end-to-end CI/CD pipelines, including creating a resilient artifact store to manage all model checkpoints across multiple versions and providers.
- Develop and maintain user-friendly APIs and interaction patterns that empower our product and research teams to ship groundbreaking features at high velocity.
- Manage and optimize our complex inference workloads at scale, operating across multiple clusters and hardware providers.
- 5+ years of professional engineering experience with deep, hands-on proficiency in Python and complex distributed systems architecture.
- Extensive, practical experience building and managing systems at scale, specifically with queues, scheduling, traffic-control, and fleet management.
- Deep expertise in our core infrastructure stack: Linux, Docker, and Kubernetes.
- Strong experience with Redis, S3-compatible storage, and public cloud platforms (AWS).
- Experience with high-performance, large-scale ML systems (managing >100 GPUs).
- Deep familiarity with PyTorch and CUDA.
- Experience with modern networking stacks, including RDMA (RoCE, Infiniband, NVLink).
- Familiarity with FFmpeg and multimedia processing pipelines.
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say
