Elastix AI

AI Software Engineer

Seattle, WA
Python C++ Docker Kubernetes PyTorch
Description

AI Software Engineer

Department: Engineering

Location: Seattle

Employment Type: FullTime

About Elastix AI

We are building the next-gen AI inference platform.

Description

Job Title: Software Engineer, AI Inference Platform

Company: ElastixAI, Inc.

Location: Seattle, WA (Hybrid - 3 days/week in office)

About ElastixAI

ElastixAI is an early-stage startup building the next-generation AI inference infrastructure — co-designed across ML software and custom accelerator hardware. Our platform dynamically optimizes inference efficiency and scalability across diverse deployments, enabling adaptive, high-performance AI serving.

Role Summary

We’re looking for a systems-minded AI Software Engineer to join our core inference platform team. You’ll design and extend the low-level serving stack — hacking open-source frameworks like vLLM, SGLang, and TensorRT-LLM, building new model sharding and scheduling logic, and integrating deeply with our proprietary AI accelerator. This role sits at the intersection of ML systems, compiler/runtime engineering, and hardware-software co-design.

Key Responsibilities

  • Architect, extend, and optimize core components of our AI serving platform for throughput, latency, and scalability.

  • Customize open-source serving frameworks (e.g., vLLM) for proprietary model ingestion and accelerator integration.

  • Develop efficient model partitioning, scheduling, and memory management strategies for multi-device inference.

  • Collaborate with ML engineers on model export and runtime optimization (quantization, graph transforms).

  • Work closely with hardware engineers to influence accelerator interface design and performance tuning.

  • Build APIs and runtime tools enabling flexible, PyTorch-native model deployment on our infrastructure.

  • Profile, debug, and optimize across the full stack — from Python orchestration to C++ kernels and PCIe drivers.

Required Qualifications

  • BS/MS/PhD in Computer Science, Electrical/Computer Engineering, or related field.

  • 3+ years of professional experience in systems programming, ML infrastructure, or distributed inference.

  • Proficient in C++ and Python, with strong debugging and performance analysis skills.

  • Deep familiarity with one or more LLM serving frameworks (vLLM, SGLang, TensorRT-LLM, DeepSpeed-Inference, etc.).

  • Understanding of model deployment internals — token scheduling, KV caching, batching, and pipelined inference.

  • Comfortable working close to the hardware abstraction layer — CUDA, PCIe, memory management, or runtime scheduling.

  • Strong collaboration and communication skills; ability to work cross-functionally in a fast-paced startup environment.

Preferred / Bonus

  • Experience with hardware-aware ML optimization, compiler/runtime integration, or accelerator SDKs.

  • Hands-on experience profiling GPU/accelerator workloads.

  • Familiarity with containerized deployments (Docker/Kubernetes).

  • Exposure to distributed systems or large-scale inference clusters.

  • Contributions to open-source ML or serving frameworks.

What We Offer:

  • A chance to be a foundational engineer in an innovative AI startup

  • A dynamic and collaborative work environment and the change to have a significant impact on new technology

  • The opportunity to work on challenging problems at the intersection of ML, software, and systems.

  • Competitive compensation and startup equity package

  • Comprehensive medical, dental, and vision coverage (100% paid by employer)

  • Life insurance and AD&D

  • Flexible Time Off (FTO)

  • 12-paid holidays

  • Paid parental leave

  • Gym or fitness benefit

  • Commuter benefit

  • Weekly catered lunches in the office

  • Investment in employee learning & development

Elastix AI
Elastix AI

0 applies

0 views

There are more than 50,000 engineering jobs:

Subscribe to membership and unlock all jobs

Engineering Jobs

60,000+ jobs from 4,500+ well-funded companies

Updated Daily

New jobs are added every day as companies post them

Refined Search

Use filters like skill, location, etc to narrow results

Become a member

🥳🥳🥳 452 happy customers and counting...

Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.

To try it out

For active job seekers

For those who are passive looking

Cancel anytime

Frequently Asked Questions

  • We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
  • We've got over 200,000 jobs from 15,000+ vetted companies. No fake or sleazy jobs here!
  • We aggregate jobs from 15,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
  • We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
  • Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
  • Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
  • Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅

What Fellow Engineers Say