Member of Technical Staff, Model Evaluation
Location: Bay Area
Department: Research
Location Type: IN_OFFICE
Employment Type: FULL_TIME
- Design, develop, and maintain robust evaluation frameworks and benchmarks for measuring LLM performance across diverse tasks and domains.
- Define and implement quantitative metrics that capture model quality, safety, reliability, and regression detection.
- Build scalable, automated evaluation pipelines that integrate into model training and deployment workflows.
- Conduct rigorous statistical analysis of model outputs to identify failure modes, biases, and performance gaps.
- Partner with product and customer-facing teams to translate real-world use cases into meaningful evaluation criteria.
- BS/MS/PhD in Computer Science, Machine Learning, Statistics, or a related field (or equivalent experience).
- At least 2 years of experience in ML evaluation, applied ML research, or a related engineering role.
- Strong understanding of LLM fundamentals (autoregressive generation, instruction tuning, RLHF, in-context learning, decoding strategies).
- Proficiency in Python and ML frameworks such as PyTorch.
- Experience designing and implementing evaluation metrics and benchmarks for generative models.
- Solid foundation in statistics, experimental design, and hypothesis testing.
- Experience with version control (Git) and containerization (Docker).
- Excellent communication skills with the ability to distill complex evaluation results into actionable insights.
- Experience with human-in-the-loop evaluation systems (Likert-scale annotation, pairwise preference ranking, red-teaming).
- Familiarity with LLM safety and alignment evaluation (toxicity, hallucination detection, factual grounding).
- Knowledge of existing benchmark suites (MMLU, HumanEval, HELM, BIG-Bench) and their limitations.
- Experience building evaluation infrastructure at scale using cloud platforms (AWS, GCP, Azure).
- Familiarity with MLOps practices and CI/CD pipelines for model validation.
- Experience with data engineering, large-scale data labeling, or synthetic data generation for evaluation purposes.
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say
