Majestic Labs logo

LLM Inference Engineer

Majestic Labs

On-site
Los Altos, CA
Full-time
Senior
3+ yrs
Salary not listedPosted 2mo ago

Real job — pulled straight from Majestic Labs’s careers page · Verified June 16, 2026 · No reposts.

Job description

Majestic Labs is hiring a LLM Inference Engineer — a full-time, based in Los Altos, CA role. Apply directly on Majestic Labs's careers page below.

LLM Inference Engineer

Location: Los Altos, California (US)

Experience Level: Senior

Description

About Majestic Labs

We’re a fast-moving, US-Israeli AI startup building next-generation infrastructure for the world’s most demanding AI workloads. Our mission is to accelerate the future of intelligence and make it ubiquitously accessible by delivering full stack custom AI servers optimized for large-scale inference and training.

Backed by top-tier investors and led by industry veterans, we’re scaling rapidly. We seek forward-thinking engineers and operators who thrive in a collaborative environment and excel at solving complex challenges from first principles. If you want to build infrastructure that fundamentally changes what the world can accomplish with AI, come join us!

About the position

In this high-impact role, you are the bridge between cutting-edge custom silicon and production-grade AI. You will own the end-to-end LLM serving stack on Majestic hardware, architecting everything from serving APIs down to KV cache management, batching, and scheduling. Your primary mission is to port leading frameworks like vLLM and SGLang to our accelerator and optimize them for peak performance. Because our architecture offers memory headroom, you won't just match traditional GPUs; you will shatter their limits on throughput, batch sizes, and context lengths. As you hunt down bottlenecks, your insights will directly steer our future kernel, compiler, and hardware development. 

Responsibilities

  • The serving stack, end to end — bring up and adapt a modern inference framework (vLLM, SGLang, or similar) to run on Majestic hardware.
  • The runtime hot path — continuous batching, the scheduler, paged KV cache, and prefill/decode disaggregation.
  • Distributed inference at scale — tensor, pipeline, and expert parallelism across accelerators, wired into our collective communication library (CCL).
  • The multi-modal pipeline — image, audio, and video preprocessing, encoder integration, and mixed-modality batching.
  • Inference-time techniques — speculative decoding, prefix caching, and structured decoding.
  • End-to-end performance — profile, benchmark, and hunt down bottlenecks across the full serving path, feeding findings back to the kernel, compiler, and hardware teams.

Requirements

Requirements

  • 3+ years building or operating production LLM inference and serving systems (5+ preferred).
  • Deep, hands-on work with a modern inference framework vLLM, SGLang, TensorRT-LLM, Fireworks, or similar including its scheduler, paged attention / KV cache, model executor, and backend integration points.
  • Strong Python and C++, with the ability to move fluidly between the two.
  • A real grasp of transformer inference the prefill/decode split, KV cache behavior, and how batching dynamics shape latency and throughput.
  • Distributed inference experience tensor and pipeline parallelism across multiple devices.
  • An instinct for performance you can profile an end-to-end stack and chase a regression from the serving API all the way down to the kernel.

At Majestic Labs, we are shaping the next paradigm and making AI ubiquitous by building radically better infrastructure across the entire stack-from systems engineering and VLSI to compilers and AI applications. This epic mission requires a relentless culture of innovation, grit, and dedication. Our employees are our most important asset.

Have you read the job description and feel you could be a great fit?

We seek smart, curious, self-driven individuals who enjoy creative problem-solving and continuous learning. Even if you don't meet all requirements, we'd love to hear from you if you have the mindset to tackle complex challenges.

Majestic Labs is an equal opportunity employer deeply committed to diversity of background and thought. If you are ready to dare greatly, learn from mistakes, engage in vigorous debates, and help create a step-function shift in human technical capability-we want to meet you!

Get LLM Inference Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Sequoia Connect logo

Senior Azure Developer

Remote · Mexico-eligible
✓ From careers page· 42m ago
Sequoia Connect logo

Senior Data Engineer (Remote)

Remote · US-eligible
✓ From careers page· 42m ago
Sequoia Connect logo

Senior AWS Data Engineer

United States
✓ From careers page· 42m ago
Sequoia Connect logo

Senior AWS Data Engineer

México
✓ From careers page· 42m ago

Frequently asked questions

What skills are required for LLM Inference Engineer at Majestic Labs?

The required skills for LLM Inference Engineer at Majestic Labs include: Python, C++.

What is the seniority level for LLM Inference Engineer at Majestic Labs?

LLM Inference Engineer at Majestic Labs is a Senior level position.

How do I apply for LLM Inference Engineer at Majestic Labs?

You can view the full description and apply for LLM Inference Engineer at Majestic Labs on EchoJobs: https://echojobs.io/job/majestic-labs-llm-inference-engineer-mqdlr.