Member of Technical Staff, Reinforcement Learning
Location: Bay Area
Department: Research
Location Type: IN_OFFICE
Employment Type: FULL_TIME
- Design, develop, and optimize RL training pipelines (PPO, DPO, RLHF, and novel approaches) for diffusion-based LLMs.
- Build and iterate on reward models, reward shaping strategies, and evaluation of reward quality.
- Implement innovative approaches for fine-tuning and scaling generative AI models.
- Work on data preprocessing pipelines, model evaluation, and alignment to enterprise use cases.
- Research and implement techniques for controlled text generation and constraint satisfaction.
- Improve training stability, efficiency, and reproducibility of RL workloads.
- BS/MS/PhD in Computer Science or a related field (or equivalent experience).
- At least 2 years of experience working on ML projects in PyTorch (or equivalent), preferably in a research lab or engineering role.
- Excellent familiarity with transformers and core LLM concepts (autoregressive pretraining, instruction tuning, in-context learning, KV caching).
- Hands-on experience with reinforcement learning from human feedback (RLHF), PPO, DPO, or related post-training methods.
- Familiarity with training and inference in diffusion models.
- Experience training deep learning models at scale in distributed computing environments.
- Extensive experience training transformer-based language models from scratch.
- Experience designing and implementing reward models or preference learning systems.
- Knowledge of advanced training techniques (mixed precision, gradient accumulation, etc.).
- Background in optimization theory and neural network architecture design.
- Experience with LLM serving frameworks like vLLM, SGLang, or TensorRT.
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say
