
Inference Optimization Intern, Performance Modeling
Institute for Foundation Models
Real job — pulled straight from Institute for Foundation Models’s careers page · Verified July 13, 2026 · No reposts.
Job description
Institute for Foundation Models is hiring a Inference Optimization Intern, Performance Modeling — a internship, based in Sunnyvale, CA role. Apply directly on Institute for Foundation Models's careers page below.
Inference Optimization Intern – Performance Modeling
Team: Engineering
Location: Sunnyvale, CA
Commitment: Intern | Fall
Workplace Type: onsite
Key Responsibilities
-
Develop analytical performance models for GPU kernels and inference workloads.
-
Build and validate a simulator to estimate theoretical hardware performance limits.
-
Compare measured kernel performance against architectural peak throughput.
-
Identify performance bottlenecks in compute, memory, communication, and scheduling.
-
Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
-
Investigate PTX and SASS code generation to understand low-level execution behavior.
-
Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
-
Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
-
Design profiling methodologies for Hopper and Blackwell architectures.
-
Document findings and provide actionable recommendations for performance improvements.
Academic Qualifications
Preferred Qualifications
-
Experience with CUDA programming and GPU kernel development.
-
Understanding of NVIDIA GPU architecture and memory hierarchy.
-
Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
-
Knowledge of PTX, SASS, and low-level GPU execution.
-
Experience optimizing CUDA kernels for throughput and latency.
-
Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
-
Experience with deep learning frameworks such as PyTorch or TensorFlow.
-
Strong programming skills in C++, CUDA, and Python.
Desired Skills
-
Performance engineering mindset.
-
Strong analytical and debugging abilities.
-
Interest in AI systems, inference optimization, and hardware-software co-design.
-
Ability to work independently on research and engineering challenges.
-
Excellent written and verbal communication skills.
Get Inference Optimization Intern, Performance Modeling jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Inference Optimization Intern, Performance Modeling at Institute for Foundation Models?
The required skills for Inference Optimization Intern, Performance Modeling at Institute for Foundation Models include: PyTorch, TensorFlow, C++, Python.
What is the seniority level for Inference Optimization Intern, Performance Modeling at Institute for Foundation Models?
Inference Optimization Intern, Performance Modeling at Institute for Foundation Models is a Internship level position.
How do I apply for Inference Optimization Intern, Performance Modeling at Institute for Foundation Models?
You can view the full description and apply for Inference Optimization Intern, Performance Modeling at Institute for Foundation Models on EchoJobs: https://echojobs.io/job/institute-for-foundation-models-inference-optimization-intern-performance-modeling-y3oa4.