
Real job — pulled straight from Modular’s careers page · Verified August 21, 2026 · No reposts.
Job description
Modular is hiring a Developer Advocate, MAX Inference & Serving (Remote) — a full-time, remote role ($150k–$225k). Apply directly on Modular's careers page below.
Developer Advocate, MAX Inference & Serving
Location: United States / Canada
Department: Developer Relations
Location Type: REMOTE
Employment Type: FULL_TIME
About the role:
What you will do:
- Build and publish reproducible benchmarks comparing MAX against vLLM, SGLang, TensorRT-LLM, and Triton Inference Server, including methodology, harness code, and hardware configurations so others can verify the numbers.
- Deploy and profile real inference workloads on MAX across CPU and GPU targets, and investigate performance gaps in latency, throughput, and cost per token.
- Write and maintain technical content grounded in that work: performance deep dives, serving architecture explainers, and posts that show how MAX handles batching, KV cache management, quantization, and multi-GPU serving.
- Author runnable tutorials, examples, and video walkthroughs covering model deployment on MAX, from
pip install modularto a served endpoint under load. - Provide technical support to engineers evaluating MAX for inference and serving, including teams at some of the world's largest companies, and debug their deployment and performance issues directly.
- Answer technical questions across GitHub, Discord, X, and LinkedIn, reproducing reported issues and filing them with enough detail for engineering to act on.
- Translate community and customer findings into prioritized inference and serving feedback for engineering and product, shaping the MAX roadmap.
- Present technical talks, benchmark results, and live demos at conferences, summits, and meetups.
- Help set the technical direction of Modular's developer relations content, including what gets benchmarked, documented, and demoed next.
What you bring to the table:
- 3-5 years of professional engineering experience, with at least some of it spent deploying or operating ML systems in production. You have run inference workloads yourself, not just written about them.
- Deep familiarity with the ML inference stack. You can explain continuous batching, KV cache management, quantization trade-offs, and tensor and pipeline parallelism, and you have hands-on experience with at least one of vLLM, Triton Inference Server, TensorRT-LLM, or SGLang.
- Strong Python and systems programming experience in C++, Rust, Mojo, or CUDA is a significant advantage, especially if you have profiled and optimized GPU code.
- Comfort with the deployment surface around serving: containers, Kubernetes, GPU drivers and runtimes, and the usual ways a cluster refuses to cooperate.
- You benchmark rigorously. You know why one number is not a result, how to control for warmup and batch size, and when a comparison is not apples to apples.
- You write well about technical work, and you have a portfolio to show it: blog posts, tutorials, videos, docs, or courses. You explain complex ideas without losing precision, and you cut the filler.
- You learn new tools fast and turn them into accurate content quickly. A feature ships Tuesday, your tutorial goes out Thursday and the code runs.
- A growth and leadership mindset, with a collaborative attitude that seeks to learn more from our customers, team members, and the broader market
Minimum Qualifications:
- Bachelor's degree in Engineering, Information Systems, Computer Science, or technical related field.
- 2+ years of Product Management or related work experience.
Helpful, but not required:
- Experience programming GPUs using CUDA or ROCm.
- Familiarity with the Mojo 🔥 programming language and MAX AI framework.
- Familiarity with open source software development practices and communities.
- Experience producing high-quality videos covering technical topics.
What Modular brings to the table:
- Amazing Team. We are a progressive and agile team with some of the industry’s best engineering and product leaders.
- World-class Benefits. In order to attract the best, we need to offer the best. Your benefits package may include comprehensive healthcare coverage, retirement and savings programs, employee stock purchase opportunities, paid time off, wellbeing resources, family support programs, and learning and development opportunities. Please note that specific benefit packages may vary based on your location, you can read more about benefits offered by Qualcomm here.
- Competitive Compensation. We offer very strong compensation packages, including RSU grants. We want people to be focused on their best work and believe in tailoring compensation plans to meet the needs of our workforce.
- Team Building Events. We organize regular team onsites and local meetups in Los Altos, CA as well as different cities. Traveling 2-4 times a year is expected for all roles.
Get Developer Advocate, MAX Inference & Serving (Remote) jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs

Senior System Test Engineer, Machine Safety & Industrial Communications



Senior Machine Learning Engineer, Private Search Quality (Remote)
Frequently asked questions
What is the salary for Developer Advocate, MAX Inference & Serving (Remote) at Modular?
The estimated salary range for Developer Advocate, MAX Inference & Serving (Remote) at Modular is $150,000 - $225,000 USD per year.
Is Developer Advocate, MAX Inference & Serving (Remote) at Modular a remote job?
Yes, Developer Advocate, MAX Inference & Serving (Remote) at Modular is a remote position. Candidates in Los Altos, CA may be preferred.
What skills are required for Developer Advocate, MAX Inference & Serving (Remote) at Modular?
The required skills for Developer Advocate, MAX Inference & Serving (Remote) at Modular include: Python, C++, Rust, Machine Learning, GitHub.
What is the seniority level for Developer Advocate, MAX Inference & Serving (Remote) at Modular?
Developer Advocate, MAX Inference & Serving (Remote) at Modular is a Mid Level / Senior level position.
How do I apply for Developer Advocate, MAX Inference & Serving (Remote) at Modular?
You can view the full description and apply for Developer Advocate, MAX Inference & Serving (Remote) at Modular on EchoJobs: https://echojobs.io/job/modular-developer-advocate-max-inference-serving-76pnq.