
Real job — pulled straight from Dragonfly’s careers page · Verified August 14, 2026 · No reposts.
Job description
Dragonfly is hiring a Senior Inference Optimization Engineer — a full-time, remote role. Apply directly on Dragonfly's careers page below.
Senior Inference Optimization Engineer - Dragonfly Portfolio
Location: United States (Remote)
Department: Portfolio
Location Type: HYBRID
Employment Type: FULL_TIME
- 5+ years in performance optimization or HPC with deep GPU architecture and parallel programming knowledge
- Hands-on experience with at least one production LLM inference engine (vLLM, SGLang) running at high volume
- Demonstrated experience with LLM inference optimization: continuous batching, PagedAttention, KV cache management, speculative decoding, quantization, CUDA graphs, torch.compile
- Experience with distributed inference strategies: tensor parallelism, pipeline parallelism, MoE parallelism in multi-GPU and multi-node environments
- GPU profiling fluency: Nsight Systems, Nsight Compute, PyTorch Profiler
- Proficiency in Python, Rust, or Go. C++/CUDA a strong plus
- Bonus: custom Triton kernels, diffusion/image model inference optimization, open-source inference framework contributions
- Stand up and optimize GPU infrastructure including B300 nodes in owned data centers
- Drive down TTFT and TPOT, push throughput, and improve cost per token for LLM inference workloads
- Build reproducible benchmarking harnesses across inference engines to identify optimal engine, quantization scheme, and parallelism strategy per workload and GPU SKU
- Optimize multivariate inference load-balancing algorithms within the inference routing system
- Evaluate emerging inference optimization techniques including custom CUDA/Triton kernels, novel attention variants, new quantization schemes, and compilation stack improvements
- Evaluate emerging inference hardware (FPGAs, ASICs, custom silicon) for viability in the stack
- We'll review your application and assess fit for this role.
- If there's a match, we'll facilitate a warm introduction to the team.
- If the timing isn't right, we'll keep you in mind for future opportunities across the portfolio.
Get Senior Inference Optimization Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs



Senior Machine Learning Software Verification Engineer

Frequently asked questions
Is Senior Inference Optimization Engineer at Dragonfly a remote job?
Yes, Senior Inference Optimization Engineer at Dragonfly is a remote position. This role is open to remote candidates.
What skills are required for Senior Inference Optimization Engineer at Dragonfly?
The required skills for Senior Inference Optimization Engineer at Dragonfly include: Python, Rust, Go, C++, LLM, Deep Learning.
What is the seniority level for Senior Inference Optimization Engineer at Dragonfly?
Senior Inference Optimization Engineer at Dragonfly is a Senior level position.
How do I apply for Senior Inference Optimization Engineer at Dragonfly?
You can view the full description and apply for Senior Inference Optimization Engineer at Dragonfly on EchoJobs: https://echojobs.io/job/dragonfly-senior-inference-optimization-engineer-dragonfly-portfolio-6kgso.