Do you want to be part of the team that brings Artificial Intelligence (AI) emerging technology to the field? We are looking for a hardworking Solution Architect (SA) to join the NVIDIA AI Enterprise (NVAIE) SA Segment Team. The mission of the NVAIE Segment team is to guide and enable the successful adoption at scale of NVIDIA AI Enterprise Software in production.
In our Solutions Architecture team, we work with NVIDIA's pioneering hardware and software, driving the latest breakthroughs in artificial intelligence. We need people who enable customer adoption of NVIDIA technology and develop lasting relationships with our technology partners, making NVIDIA a key design choice for end-user solutions. On this team, you will support full stack deployment including architectural designs, workload orchestration and application optimization. At NVIDIA, you will be immersed in a diverse, encouraging environment where everyone is inspired to do their life's work. Come join the team and see how you can make a lasting impact on the world!
What You’ll Be Doing:
Primary responsibilities will include building and enabling robust AI/HPC infrastructure for customers
Support operational and reliability aspects of large-scale AI clusters, focusing on performance at scale, training stability, real-time monitoring, logging, and alerting
Engage in and improve services from inception and design through deployment, operation, and optimization
Co-design telemetry of AI workloads to help engineering build solutions for more robust workloads at scale
Communicate across internal teams to support the continuous improvement of NVIDIA's offerings and software designs
What We Need to See:
Strong foundational expertise, from a BS, MS, or Ph.D. degree in Engineering, Mathematics, Physics, Computer Science, Data Science, or similar (or equivalent experience).
8+ years of experience and knowledge of neural networks including good understanding of transformer architectures. Experience designing large scale AI workloads with SLURM and/or Kubernetes
Proficiency with Python / C++ / Rust or other popular software languages
Excellent verbal, written communication, and technical presentation skills in English
You are motivated to work with multiple levels and teams across organizations
Strong analytical and problem-solving skills
Strong time-management and organization skills for coordinating multiple initiatives, priorities and implementations of new technology and products into very sophisticated projects
You are a curious self-starter with a desire for continuous learning and sharing knowledge across the team
Ways to Stand Out from The Crowd:
Experience orchestrating distributed Deep Learning training with SLURM
Proficiency in DevOps, including hands-on experience with Ansible, Terraform or similar tools. Equivalent experience will be accepted as well.
8+ years designing solutions with one or more Tier-1 Clouds (AWS, Azure, GCP or OCI) and cloud-native architectures and software
Technical leadership with a strong understanding of NVIDIA technologies, and success in working with customers
Expertise with parallel file systems (e.g. Lustre, GPFS, BeeGFS, WekaIO) and high-speed interconnects (InfiniBand, Omni Path, and Gig-E)
Experience with integration and deployment of software products in production enterprise environments, and microservices software architecture
You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.
Other Jobs from NVIDIA
Senior CPU Verification Engineer
Senior Verification Engineer - Memory Subsystem
Senior Verification Engineer - Memory Subsystem
Design Verification Infrastructure Engineer
Chip Architect Engineer
Similar Jobs
Staff Machine Learning Engineer, Data Infrastructure, Central
There are more than 50,000 engineering jobs:
Subscribe to membership and unlock all jobs
Engineering Jobs
60,000+ jobs from 4,500+ well-funded companies
Updated Daily
New jobs are added every day as companies post them
Refined Search
Use filters like skill, location, etc to narrow results
Become a member
🥳🥳🥳 452 happy customers and counting...
Overall, over 80% of customers chose to renew their subscriptions after the initial sign-up.
To try it out
For active job seekers
For those who are passive looking
Cancel anytime
Frequently Asked Questions
- We prioritize job seekers as our customers, unlike bigger job sites, by charging a small fee to provide them with curated access to the best companies and up-to-date jobs. This focus allows us to deliver a more personalized and effective job search experience.
- We've got about 70,000 jobs from 5,000 vetted companies. No fake or sleazy jobs here!
- We aggregate jobs from 5,000+ companies' career pages, so you can be sure that you're getting the most up-to-date and relevant jobs.
- We're the only job board *for* software engineers, *by* software engineers… in case you needed a reminder! We add thousands of new jobs daily and offer powerful search filters just for you. 🛠️
- Every single hour! We add 2,000-3,000 new jobs daily, so you'll always have fresh opportunities. 🚀
- Typically, job searches take 3-6 months. EchoJobs helps you spend more time applying and less time hunting. 🎯
- Check daily! We're always updating with new jobs. Set up job alerts for even quicker access. 📅
What Fellow Engineers Say