
Software Engineer, Machine Learning Infrastructure
Real job — pulled straight from realm-labs’s careers page · Verified June 5, 2026 · No reposts.
Job description
realm-labs is hiring a Software Engineer, Machine Learning Infrastructure — a full-time, based in Sunnyvale, CA role. Apply directly on realm-labs's careers page below.
Software Engineer, ML Infrastructure
Location: Sunnyvale, CA
Department: Engineering
Location Type: IN_OFFICE
Employment Type: FULL_TIME
Role Overview
About Realm Labs
Key Responsibilities
- Own the end-to-end LLM inference stack, including:
- Model loading and execution
- GPU utilization and memory efficiency
- Runtime performance tuning
- Production deployment and scaling
- Design and operate high-performance LLM serving systems using technologies such as:
- vLLM, TensorRT / TensorRT-LLM, Triton Inference Server, SGLang
- Optimize inference across:
- Latency
- Throughput (QPS)
- GPU memory footprint
- Cost efficiency
- Work hands-on with PyTorch and TensorFlow models, including:
- Model graph understanding
- Attention mechanisms, KV cache behavior, batching strategies
- Precision tradeoffs (FP16, BF16, INT8, etc.)
- Build and maintain production-grade GPU services:
- Multi-model serving
- Autoscaling strategies
- Fault isolation and graceful degradation
- Collaborate with application and platform teams to:
- Define serving APIs
- Ensure correctness and safety of outputs
- Debug production issues end-to-end
- Build a reproducible model training and versioning system for customer deployments
- Establish best practices for:
- Model versioning
- Rollouts and rollbacks
- Performance benchmarking
- Production validation
Expected Qualifications
- 5+ years of professional experience in ML infrastructure, systems engineering, or production ML roles.
- Strong software engineering fundamentals; ability to write robust, maintainable production code.
- Deep hands-on experience with LLM inference infrastructure, including:
- PyTorch (required)
- TensorFlow (working knowledge)
- Proven experience with GPU inference optimization, including:
- TensorRT / TensorRT-LLM
- vLLM
- Triton Inference Server
- SGLang or similar serving runtimes
- Strong understanding of LLM internals, such as:
- Transformer architectures
- Attention and KV caching
- Batching, streaming, and token-level generation
- Experience running ML systems in production with high traffic and SLAs
- Comfortable working in Linux-based, cloud production environments
Preferred Qualifications
- Experience deploying LLMs on Kubernetes and GPU clusters.
- Familiarity with CUDA, NCCL, or low-level GPU performance concepts.
- Experience with:
- Model sharding and parallelism strategies
- Multi-GPU inference
- Streaming inference systems
- Knowledge of observability for ML systems (metrics, latency breakdowns, GPU monitoring).
- Experience working at startups or owning systems with minimal abstraction layers.
Additional Information
- This is a founding, high-ownership role with direct impact on core product capabilities.
- You will be expected to build, run, and own systems end-to-end.
- The role may include limited on-call responsibilities aligned with production ownership.
Compensation & Benefits
- Market aligned compensation and benefits
- Founding engineer equity (Equity is a significant component of this role and will be discussed)
- Medical, Dental, Vision, Life insurance, 401-K, In-office lunch etc.
Get Software Engineer jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Software Engineer, Machine Learning Infrastructure at realm-labs?
The required skills for Software Engineer, Machine Learning Infrastructure at realm-labs include: Python, TensorFlow, PyTorch, Kubernetes, Linux, LLM.
What is the seniority level for Software Engineer, Machine Learning Infrastructure at realm-labs?
Software Engineer, Machine Learning Infrastructure at realm-labs is a Senior level position.
How do I apply for Software Engineer, Machine Learning Infrastructure at realm-labs?
You can view the full description and apply for Software Engineer, Machine Learning Infrastructure at realm-labs on EchoJobs: https://echojobs.io/job/realm-labs-software-engineer-ml-infrastructure-biz5m.