
Senior AI Inference Engineer, Model Optimization & Deployment
Real job — pulled straight from Zoox’s careers page · Verified July 16, 2026 · No reposts.
Job description
Zoox is hiring a Senior AI Inference Engineer, Model Optimization & Deployment — a full-time, based in Foster City, CA role. Apply directly on Zoox's careers page below.
Senior AI Inference Engineer - Model Optimization & Deployment
Team: Autonomy Software
Location: Foster City, CA, San Diego, CA, Seattle, WA
Commitment: Full-time
Workplace Type: hybrid
Salary:
As a Model Optimization & Deployment Engineer, you will focus on bringing highly efficient, production-ready large-scale models to our on-vehicle stack. We are looking for experts with hands-on experience in compressing, accelerating, and deploying complex models (LLMs, VLMs, or FMs) for power- and thermal-constrained vehicle SOCs. You will optimize the ML models, write custom CUDA kernels, and build highly concurrent inference code to ensure real-time, deterministic execution on edge devices.
In this role, you will:
- Optimize large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs) using advanced quantization (PTQ, QAT), pruning, mixed-precision inference frameworks, and parameter-efficient fine-tuning (LoRA, QLoRA).
- Architect and implement model conversion and compilation pipelines using TensorRT for edge deployment.
- Perform rigorous parity checking, accuracy recovery, and latency benchmarking between PyTorch frameworks and compiled edge binaries.
- Develop and optimize custom ML OPs and TensorRT Plugins with efficient CUDA kernels to minimize latency and maximize memory bandwidth on AI accelerators.
- Write production-level, low latency, and memory-safe C++ and CUDA code for real-time inference on vehicle systems.
Qualifications:
- Deep expertise in model quantization (PTQ, QAT) and mixed-precision inference frameworks (INT8, FP8, FP4, BF16/FP16).
- Proven experience optimizing large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs/VLAs) utilizing Efficient Attention mechanisms (e.g., FlashAttention, Linear Attention), KV-cache optimization (e.g., PagedAttention) and Speculative Decoding.
- Extensive experience with model conversion/compilation pipelines (e.g., ONNX, TensorRT, torch.compile) and performing rigorous latency benchmark and model quality parity valuation.
- Proficiency in low-level programming for AI accelerators, specifically developing and optimizing custom ML OPs and TensorRT Plugins with efficient CUDA kernel implementations.
- Production-level C++ (14/17/20) and Python programming skills, with experience developing concurrent, memory-safe, real-time inference code for edge devices.
Bonus Qualifications:
- Familiarity with SOTA autonomous driving perception algorithms (temporal 3D object detection, BEV, 3D Occupancy Networks) and multi-modal sensor processing (Vision, LiDAR, Radar).
- Experience with distributed training pipelines and model/tensor parallelism (PyTorch Distributed, Ray, DeepSpeed, Megatron-LM) and runtime efficiency optimization for GPU clusters.
- Experience with end-to-end autonomous driving paradigms (VLM/VLA models, Foundation models) and edge deployment technologies (e.g., TensorRT-LLM).
Get Senior AI Inference Engineer, Model Optimization & Deployment jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Senior AI Inference Engineer, Model Optimization & Deployment at Zoox?
The required skills for Senior AI Inference Engineer, Model Optimization & Deployment at Zoox include: Python, C++, PyTorch, LLM, Machine Learning.
What is the seniority level for Senior AI Inference Engineer, Model Optimization & Deployment at Zoox?
Senior AI Inference Engineer, Model Optimization & Deployment at Zoox is a Senior level position.
How do I apply for Senior AI Inference Engineer, Model Optimization & Deployment at Zoox?
You can view the full description and apply for Senior AI Inference Engineer, Model Optimization & Deployment at Zoox on EchoJobs: https://echojobs.io/job/zoox-senior-ai-inference-engineer-model-optimization-deployment-rxxan.