Eka Care logo

Senior Data Scientist, Kernel Optimization & Inference Engineer

Eka Care

On-site
Bengaluru, India
Full-time
Senior
2+ yrs
Salary not listedPosted 2d ago

Real job — pulled straight from Eka Care’s careers page · Verified August 11, 2026 · No reposts.

Job description

Eka Care is hiring a Senior Data Scientist, Kernel Optimization & Inference Engineer — a full-time, based in Bengaluru, India role. Apply directly on Eka Care's careers page below.

Senior Data Scientist ( Kernel Optimisation & Inference Engineer)

Location: Bengaluru, India

Department: Data Science.

Experience: 2 - 5

Kernel Optimisation & Inference Engineer
Bengaluru · Full-time · 2–5 yrs
Somewhere between the model and the silicon, 10–20% of a training budget goes missing. Your job is to go get it back, and then make the same model fast enough to run in a clinic, and small enough to run on a phone.

About EkaCare and the mission
EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.
Now, under the IndiaAI Mission, EkaCare has been selected to build India's open healthcare foundation model: a ~30B-parameter MoE trained on ~500B+ tokens across 12+ Indic languages. Weights and the India-specific clinical eval suite ship in open domain.

The role:
You'll work with our performance lead on making everything fast: training-side fused kernels and MFU on the 30B MoE, inference-side latency and throughput, and the quantized 2B/4B on-device tier. Hardware-up: profiler first, roofline reasoning always, custom kernels when the math says so.

What you'll do
  • Profile training and inference workloads and hunt utilisation gaps across kernels, memory and comms.
  • Write and tune CUDA kernels where existing ops leave real performance on the table, and know when they don't.
  • Optimise MoE-specific paths: grouped GEMMs, all-to-all communication, expert load imbalance.
  • Build the fast inference path: vLLM-class serving, continuous batching, prompt/prefix caching for clinical-context workloads, speculative decoding.
  • Own quantisation for the 2B/4B variants (AWQ/GPTQ-class, fp8) — with eval-parity verification, not just perplexity.
  • Make on-device inference real for the hardware Indian clinics actually have.
What we look for
  • 2–5 years in GPU performance work; you've profiled real workloads and shipped optimisations with before/after numbers you can defend.
  • Working fluency in CUDA, and memory-hierarchy reasoning (coalescing, occupancy, SRAM tiling; you can explain *why* FlashAttention is fast).
  • Hands-on with a modern serving stack (vLLM, TensorRT-LLM, SGLang or similar) beyond just running it.
  • Measurement discipline: you profile before optimising and verify correctness after.
Bonus
  • fp8 experience on H100/H200-class hardware; torch.compile/inductor internals.
  • Quantisation research or on-device/mobile inference experience.
  • Open-source kernels or serving contributions.
Why this is a rare gig
  • Open source, with your name on it: weights and technical reports ship publicly.
  • India-scale mission: models for a billion people in their own languages.
  • Compute that’s rare to fine: dedicated multi-node H200 training under a national grant.
  • Small senior team: you work with the people who own the recipe.
  • A live deployment path: Government institutes, EkaCare's doctors and patients use what you ship.

Full-Time Employee Benefits
  • Medical Insurance & Accidental Insurance
  • Maternity & Paternity Benefits
  • PF, Gratuity, & Leave Encashment
  • Salary Advance Policy


Get Data Scientist jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Devoteam logo

Senior Data Scientist, Generative AI

Lyon, FR
✓ From careers page· 1h ago
PwC logo

PwC

New

Senior Associate, Agentic Automation

Bengaluru
✓ From careers page· 5h ago
Hearst logo

Full-Stack Data Engineer

Troy, MI
✓ From careers page· 7h ago
TOMRA logo

Machine Learning Engineer

Mülheim-Kärlich, RP, de
✓ From careers page· 7h ago

Frequently asked questions

What skills are required for Senior Data Scientist, Kernel Optimization & Inference Engineer at Eka Care?

The required skills for Senior Data Scientist, Kernel Optimization & Inference Engineer at Eka Care include: Deep Learning, Machine Learning, Python.

What is the seniority level for Senior Data Scientist, Kernel Optimization & Inference Engineer at Eka Care?

Senior Data Scientist, Kernel Optimization & Inference Engineer at Eka Care is a Senior level position.

How do I apply for Senior Data Scientist, Kernel Optimization & Inference Engineer at Eka Care?

You can view the full description and apply for Senior Data Scientist, Kernel Optimization & Inference Engineer at Eka Care on EchoJobs: https://echojobs.io/job/eka-care-senior-data-scientist-kernel-optimisation-inference-engineer-6aucj.