We are now seeking a Senior AI Infrastructure Engineer! NVIDIA’s Compute Architecture Group is growing our team of AI focused Infrastructure Engineers who run our internal cluster for accelerated AI and software development. As part of this team, you will help to manage a diverse cluster of GPU-accelerated systems. Your contributions will enable engineers to work efficiently with a wide variety of forward-looking hardware configurations as they vigilantly seek out opportunities for performance optimization and continuously deliver high quality software.
Our ideal candidate is versatile enough to apply expertise from many domains: system administration, performance analysis, automation, and architecture. Your work will enable the ground breaking experimentation that allows us to design the world’s most powerful systems for the most demanding computing applications. You will have a meaningful impact at a fast-moving company that is spearheading the next wave in computing technology. Join our technically diverse team of GPU architects, software engineers and infrastructure experts to unlock unprecedented performance in every domain!
What you'll be doing:
Administer an NVIDIA Internal AI cluster composed of Linux systems ranging from the world’s most powerful servers to embedded systems
Maintain the configuration of our resource management system (SLURM) to keep resource allocation efficient and aligned with organizational priorities
Automate configuration management, software updates, and maintenance of system availability using modern DevOps tools (Ansible, Gitlab, etc.)
Plan and maintain new systems that support the NVIDIA Software stack
Work directly with developers and hardware architects to debug issues, identify new requirements, and improve workflows
Actively communicate with users and management regarding resource planning and allocation
What we need to see:
5+ years of previous experience deploying and administering large scale clusters, tuned for development efforts in AI
MS in Computer Science, Computer Engineering, or EECE; or a BS (or equivalent experience).
Deep knowledge of distributed resource scheduling systems (Slurm (preferred), LSF, etc.)
Demonstrated ability to script in bash, and at least one high-level language (Python preferred)
Experience with container technologies (Docker, Singularity, etc.)
Deep understanding of operating systems, computer networks, and high-performance hardware
Ability to work well with developers, hardware architects, & test engineers
Passionate dedication to providing quality support for users
Ways to stand out from the crowd:
Prior work experience managing high performance fabrics and parallel file systems
Familiarity with CUDA and managing GPU-accelerated computing systems
Basic knowledge of deep learning frameworks and algorithms
You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.
Other Jobs from NVIDIA
Automotive DriveOS Software Architect
Director of Mechanical Engineering
Senior GPU System Performance Architect
Senior DL Algorithms Engineer - Inference Performance
Similar Jobs
Senior AI-HPC Cluster Engineer
Senior Software Engineer, DevOps (Python, SQL, AWS)
Senior Manager, Software Engineering, DevOps
Lead Software Engineer, DevOps
Senior Software Engineer, DevOps
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 401 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got about 70,000 jobs from 5,000 vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 5,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say