
Real job — pulled straight from Zoox’s careers page · Verified September 20, 2026 · No reposts.
Job description
Zoox is hiring a Manager, Model Evaluation — a full-time, based in Foster City, CA role. Apply directly on Zoox's careers page below.
Manager - Model Evaluation
Team: Autonomy Software
Location: Foster City, CA
Commitment: Full-time
Workplace Type: hybrid
Salary:
There are three major components to compensation for this position: Salary, Amazon Restricted Stock Units (RSUs), and Zoox Stock Appreciation Rights. A sign-on bonus may be offered as part of the compensation package. The listed range applies only to the base salary. Compensation will vary based on geographic location and level. Leveling, as well as positioning within a level, is determined by a range of factors, including, but not limited to, a candidate's relevant years of experience, domain knowledge, and interview performance. The salary range listed in this posting is representative of the range of levels Zoox is considering for this position.
Zoox also offers a comprehensive package of benefits, including paid time off (e.g. sick leave, vacation, bereavement), unpaid time off, Zoox Stock Appreciation Rights, Amazon RSUs, health insurance, long-term care insurance, long-term and short-term disability insurance, and life insurance.
As the Manager of Model Validation & Verification (VnV) for Behavior Autonomy, you will lead an engineering and data science team responsible for evaluating, benchmarking, and validating the machine learning models and behavioral algorithms that drive our autonomous vehicle (AV) prediction and planning stacks.
You will own the statistical frameworks, offline/online evaluation metrics, and validation pipelines that ensure our behavioral models operate safely, comfortably, and predictably. Operating at the intersection of Data Science, Machine Learning, and Safety Engineering, you will partner closely with Autonomy Software, Prediction, Planner, and ML Operations teams to establish data driven release criteria for our AV fleet.
In this role, you will:
- Team Leadership & Execution: Lead, mentor, and scale a high-performing team of Data Scientists, ML Validation Engineers, and Software Engineers while driving roadmaps, sprint execution, resource allocation, and high-throughput model releases with rigorous safety guardrails. Culture of Rigor: Foster a culture of statistical excellence, healthy skepticism, proactive risk tracking, and data-driven decision-making.
- Validation Strategy & Methodologies: Define and execute end-to-end validation strategies across offline evaluation, open/closed-loop simulation, and shadow-mode fleet benchmarking to ensure robust behavioral model performance. Statistical uncertainties, and regressions into clear, data-driven recommendations for release gating and executive leadership.
- Metrics, Release Gating & Rigor: Oversee metric development and standardization with System Safety and Autonomy teams, establishing quantitative go/no-go release criteria for Behavioral Planner and Prediction ML models while fostering statistical rigor and proactive risk management.
- Cross-Functional & Infrastructure Partnership: Partner closely with Planner, Prediction, MLOps, and Developer Efficiency teams to translate behavioral requirements into measurable validation targets, streamline dataset and evaluation pipelines, and optimize runtime and compute costs.
- Executive Communication & Decision-Making: Translate complex model performance trade-offs, statistical uncertainty, regressions, and safety risks into clear, data-driven recommendations for release decisions and executive leadership.
Qualifications:
- Experience: Masters or PhD in CS, Robotics, Applied Statistics or a related field and 3+ years of direct engineering management experience leading Data Science, Machine Learning, or V&V engineering teams, alongside 7+ years of technical experience in robotics, autonomous systems, or AI/ML.
- Domain Knowledge: Strong background in ML model validation, behavioral evaluation frameworks, system-level performance benchmarking, and statistics.
- Software & Systems Literacy: Strong technical foundation in Python and modern data/ML platforms, with exposure to or conceptual literacy in large-scale production codebases (C++ or distributed systems). Proven ability to partner with systems software engineers, review technical architecture, and understand compute/performance trade-offs, Track record of leading teams evaluating complex robotic systems
- Technical Depth: Proven familiarity with modern C++/Python ML environments, simulation frameworks, high-throughput ML evaluation pipelines.
- Cross-Functional Leadership: Demonstrated ability to navigate complex organizational trade offs between release velocity, compute cost, and safety rigor.
Bonus Qualification:
- Experience with autonomous vehicles, robotics, or other safety-critical systems.
- Experience building large-scale simulation, model evaluation, or validation infrastructure.
- Experience with reinforcement learning, generative AI, or distributed ML systems.
Get Manager, Model Evaluation jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Manager, Model Evaluation at Zoox?
The required skills for Manager, Model Evaluation at Zoox include: Python, Machine Learning, Data Science, C++, Robotics, Generative AI.
What is the seniority level for Manager, Model Evaluation at Zoox?
Manager, Model Evaluation at Zoox is a Manager level position.
How do I apply for Manager, Model Evaluation at Zoox?
You can view the full description and apply for Manager, Model Evaluation at Zoox on EchoJobs: https://echojobs.io/job/zoox-manager-model-evaluation-em5p2.