Kog AI logo

GPU Engineer

Kog AI

Hybrid
Paris, France
Full-time
Mid Level
Salary not listedPosted 4w ago

Real job — pulled straight from Kog AI’s careers page · Verified July 15, 2026 · No reposts.

Job description

Kog AI is hiring a GPU Engineer — a full-time, based in Paris, France role. Apply directly on Kog AI's careers page below.

GPU Engineer

Department: Engineering

Location: Paris, France

Employment Type: FullTime

About Kog

Kog builds the fastest LLM inference engine on standard datacenter GPUs. Our Kog Inference Engine generates 3,000 output tokens per second per request on a single 8× AMD MI300X node and 2,100 on an 8× NVIDIA H200 node (FP16, batch size 1, no speculative decoding).

The hot path is a monokernel implemented with handwritten CUDA (with PTX inline assembly) on NVIDIA, and HIP (with CDNA ISA inline assembly) on AMD.

We optimize at the low level with engine/kernel/model co-design, using reverse engineering to understand and exploit the details of how the GPU hardware works at the micro level.

We are a team of 11 people, including 10 engineers and 5 PhDs.

Test it at playground.kog.ai. Read the technical details on the Kog Labs blog.

What you will work on

You will perform experiments to understand GPU internals, find creative solutions to accelerate critical computational sections used in LLM inference, and write optimized GPU kernels accordingly. Then test, profile, and optimize again.

  • Contribute to our monokernel pipeline, the single persistent GPU program that covers the full decode pass from QKV projection to LM head sampling, across AMD and NVIDIA architectures.

  • Work on low-level GPU optimization, including impossibly-fast grid synchronizations and inter-GPU collectives, and optimized GEMV, GEMM, and attention kernels across batch sizes and context lengths, with the memory-bandwidth-bound batch-1 GEMV regime as the primary target.

  • Build profiling infrastructure inside a monokernel, including custom instrumentation, device-timestamp frameworks, and per-stage analysis to translate machine behavior into concrete engineering decisions.

  • Scale the stack to third-party MoE models such as DeepSeek v4 and Qwen 3 to push generation speed on the models that matter in production today.

  • Contribute to building AI agents that will perform GPU Engineering research and kernel optimization autonomously, calibrated to hardware target and workload, starting from the inference foundations we are building now.

What we look for

  • You have written GPU kernels where performance was the central constraint. Showing the code is a requirement to move forward in the process.

  • PyTorch custom ops are an acceptable starting point if the kernels show a genuine understanding of the hardware below the framework level.

  • Stronger signals include inline PTX or CDNA ISA in public repositories, experience with latency-sensitive execution paths, understanding of why MBU matters more than MFU at batch size 1, and a background in inference engine components.

  • A top engineering school or a PhD with concrete GPU work counts, even without industry experience.

What we offer

  • Direct access to AMD and NVIDIA datacenter GPUs from day one

  • A team where creativity and technical judgment carry weight and where the people closest to the problem shape the key decisions

  • Problems that sit on the critical path of model execution speed and that directly influence what the system can become

  • A remote-friendly working model, with one mandatory week per month in our Paris office. Travel and accommodation covered by the company.

  • Compensation aligned with top technical profiles in the Paris GPU Engineering market, including equity

Get GPU Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Grab logo

Grab

New

Principal Data Scientist

Singapore, SG
✓ From careers page· 2h ago
EvolutionIQ logo

Senior Machine Learning Engineer

$190k–$225kRemote · US-eligible
✓ From careers page· 2h ago
Syngenta Group logo

HPC Lead

Bracknell, Berkshire
✓ From careers page· 3h ago
Kainos logo

Senior AI Engineer

Birmingham
✓ From careers page· 3h ago

Frequently asked questions

What skills are required for GPU Engineer at Kog AI?

The required skills for GPU Engineer at Kog AI include: PyTorch, LLM.

What is the seniority level for GPU Engineer at Kog AI?

GPU Engineer at Kog AI is a Mid Level level position.

How do I apply for GPU Engineer at Kog AI?

You can view the full description and apply for GPU Engineer at Kog AI on EchoJobs: https://echojobs.io/job/kog-ai-gpu-engineer-j8bn3.