
Real job — pulled straight from FuriosaAI’s careers page · Verified August 13, 2026 · No reposts.
Job description
FuriosaAI is hiring a Senior Software Engineer, Inference Engine — a full-time, based in Seoul role. Apply directly on FuriosaAI's careers page below.
Sr. Software Engineer - Inference Engine (Platform Software)
Department: Software
Location: Seoul HQ
Employment Type: FullTime
About the Job
Software Engineer (Inference Engine) is responsible for developing and optimizing a high-performance inference engine for Large Language Models (LLMs) and multimodal LLMs running on FuriosaAI NPUs.
In this role, you will proactively research and apply the state-of-the-art inference optimization techniques to our inference engine. You will work in close collaboration with the compiler and hardware teams to enhance the engine's performance to its full potential.
Responsibilities
Design and implement FuriosaAI’s next-generation inference engine for large and multimodal language models—comparable in capability to frameworks such as vLLM and SGLang—optimized for throughput, latency, and memory efficiency.
Design and implement advanced inference optimizations—such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling—in our production inference engine.
Design and develop capabilities for distributed and scalable inference, including prefill–decode (PD) and encode–prefill–decode (EPD) disaggregation, disaggregated speculative decoding, and hierarchical and external KV-cache storage such as HiCache and Mooncake.
Collaborate closely with the Compiler team to co-design and optimize execution for FuriosaAI NPUs, improving system-level throughput, latency, and memory utilization.
Proactively research, evaluate, and integrate state-of-the-art inference optimization techniques and key features of LLM serving frameworks into our production inference engine.
Minimum Qualifications
BS degree in Computer Science, Engineering, or a related field, with at least 3 years of relevant industry experience, or equivalent practical experience
Proficiency in Rust or C++ programming skill
Knowledge and passion of deep learning, LLM, and/or generative AI models
Excellent problem-solving and data analysis skills.
Strong communication and collaboration skills.
Preferred Qualifications
Experience in building inference serving systems for large models, encompassing batching, scheduling, caching, and load balancing.
A deep understanding of performance optimization systems.
Proficiency in C++/CUDA or Triton kernel development
Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM.
Contact
recruit@furiosa.ai
Get Software Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs
Frequently asked questions
What skills are required for Senior Software Engineer, Inference Engine at FuriosaAI?
The required skills for Senior Software Engineer, Inference Engine at FuriosaAI include: Rust, C++, Deep Learning, LLM, Generative AI.
What is the seniority level for Senior Software Engineer, Inference Engine at FuriosaAI?
Senior Software Engineer, Inference Engine at FuriosaAI is a Senior level position.
How do I apply for Senior Software Engineer, Inference Engine at FuriosaAI?
You can view the full description and apply for Senior Software Engineer, Inference Engine at FuriosaAI on EchoJobs: https://echojobs.io/job/furiosaai-sr-software-engineer-inference-engine-platform-software-dwfdo.


