
Real job — pulled straight from Cox Exponential’s careers page · Verified July 12, 2026 · No reposts.
Job description
Cox Exponential is hiring a Founding Engineer, AI Infrastructure — a full-time, based in San Francisco, CA role. Apply directly on Cox Exponential's careers page below.
Founding Engineer, AI Infra
Location: San Francisco Bay Area, CA
Department: Goaly
Location Type: HYBRID
Employment Type: FULL_TIME
About Goaly
About the Role
Key Responsibilities
- Efficiency & performance: Improve LLM training and inference efficiency through better memory utilization, optimized parallelism, and kernel-level innovations (e.g. FlashAttention, CUDA/Triton).
- Training & RL robustness: Build scalable, stable training and RL pipelines with strong reproducibility, observability, and debuggability.
- Serving & inference optimization: Design and tune high-throughput, low-latency model serving systems, including quantization, caching, and speculative decoding.
- Scalability & infrastructure: Own end-to-end training and inference infrastructure — from data ingestion and checkpointing to multi-GPU and multi-cloud orchestration.
- Production enablement: Work closely with researchers and product engineers to turn new algorithms into reliable, production-ready systems.
Requirements
- 5+ years building or operating ML infrastructure at scale, ideally supporting large language or multimodal models.
- Deep understanding of GPU architecture, distributed training frameworks (PyTorch, DeepSpeed, Megatron, Ray), and parallelism strategies.
- Hands-on experience running inference stacks (vLLM / SGLang, TGI, Triton) and optimizing them via low-level profiling.
- Strong software engineering fundamentals in Python and one of C++/Rust/Go, with clean, reliable code shipped to production.
- Working knowledge of modern data pipelines, feature stores, and vector databases used in production AI systems.
- Comfort automating infrastructure with Kubernetes, Terraform/Pulumi, and observability stacks (Prometheus, Grafana, OpenTelemetry).
Bonus Points
- Experience deploying open-source LLMs (Llama 3, Qwen, DeepSeek) or training custom foundation models.
- Contributions to ML systems tooling (compilers, kernels, inference runtimes) or open-source infrastructure projects.
- Background in reinforcement learning, evaluation harnesses, or alignment tooling that hardens production AI systems.
Get Founding Engineer, AI Infrastructure jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Founding Engineer, AI Infrastructure at Cox Exponential?
The required skills for Founding Engineer, AI Infrastructure at Cox Exponential include: Python, C++, Rust, Go, PyTorch, Kubernetes, Terraform, Prometheus, Grafana, OpenTelemetry, Machine Learning, LLM.
What is the seniority level for Founding Engineer, AI Infrastructure at Cox Exponential?
Founding Engineer, AI Infrastructure at Cox Exponential is a Staff level position.
How do I apply for Founding Engineer, AI Infrastructure at Cox Exponential?
You can view the full description and apply for Founding Engineer, AI Infrastructure at Cox Exponential on EchoJobs: https://echojobs.io/job/cox-exponential-founding-engineer-ai-infra-4c1i1.