AMD logo

System Design Engineer, AI Cluster Software Engineer

AMD

Hybrid
Santa Clara, CA
Full-time
Mid Level
$144k–$205kPosted 5d ago

Real job — pulled straight from AMD’s careers page · Verified August 8, 2026 · No reposts.

Job description

AMD is hiring a System Design Engineer, AI Cluster Software Engineer — a full-time, based in Santa Clara, CA role ($144k–$205k). Apply directly on AMD's careers page below.



ADVANCE YOUR CAREER. ADVANCE THE WORLD.

At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And we’re looking for talent who feel the same: people who want to leave the planet better than they found it, those who don’t shy away from humanity’s challenges but are determined to help solve them.

 

AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI. Whether you’re designing next-gen processors, enabling AI breakthroughs, or creating go-to-market plans, every role at AMD contributes to something bigger — technology that moves the world forward.




THE ROLE: 

This is a hands-on role for a full-stack developer to create and deliver a range of tools and applications focused on design and deployment of large-scale AI/ML clustered infrastructure. You will be working with the latest agentic tools and patterns to develop, deploy, and maintain these applications. You’ll join a growing team of multi-disciplined engineers that operates across industry verticals as subject matter experts in the AI stack and across the cluster. 

 

THE PERSON: 

  • Demonstrated use of AI coding assistants and LLM-powered developer tools: daily user of AI agents and tools
  • Professional software development experience, including substantial experience building and supporting web applications
  • Proficiency in modern frontend development using JavaScript or TypeScript and a framework such as React, Angular, or Vue
  • Experience developing backend services and APIs using a modern server-side language or framework
  • Experience deploying and operating applications using Linux, containers, CI/CD, and cloud or on-premises infrastructure
  • Experience with automated testing, source control, code review, debugging, and production support
  • Working knowledge of web application security, authentication, authorization, and secure secrets handling  

 

KEY RESPONSIBILITIES: 

  • Partner with engineering peers, domain experts in adjacent teams, and business stakeholders to understand requirements and translate them into flexible, future-proof design solutions
  • Hands on development, iteration, and maintenance of tools and applications that codify various aspects of large-scale AI cluster design stages and cluster deployment activities
  • Design and development of cohesive interface code between disparate third party tools
  • Own features from requirements and design through deployment and ongoing maintenance
  • Work in an iterative software environment, including planning and delivering work in small increments, collaborating with stakeholders, often in different areas of domain expertise (Agile development practices)
  • Participate in code reviews and retros; adapt to changing requirements and priorities

 

PREFERRED EXPERIENCE: 

  • Strong Linux fundamentals: Linux operating systems, networking, filesystems, containers, performance tooling (perf, flamegraphs, nvprof/rocprof, basic eBPF).
  • Clear communication: ability to turn complex systems into accessible, structured documentation with diagrams and reproducible steps
  • AMD ecosystem experience: ROCm, RCCL, Instinct GPUs, EPYC platforms, compiler/toolchain impacts, and performance tuning
  • Orchestration models: Slurm configuration patterns, Kubernetes for HPC/AI (GPU operators, device plugins), Apptainer/Singularity
  • Automation, IaC , and scripting tools/languages (Ansible, Terraform, Python, bash)
  • Storage/data: knowledge of or familiarity with parallel filesystems (Lustre, BeeGFS), object stores, RDMA, data pipeline throughput and caching strategies
  • Hands-on familiarity with on-premises infrastructure, particularly for AI/ML/HPC workloads would be beneficial 

 

ACADEMIC CREDENTIALS: 

  • Bachelors or Masters degree in computer science or software/computer engineering 

 

LOCATIONS:

  • Santa Clara, CA
  • Austin, TX
  • Secaucus, NJ
  • Seattle, WA 

 

This role is not eligible for visa sponsorship.

 

#LI-CB1

#LI-Hybrid




Benefits offered are described: AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.



THE ROLE: 

This is a hands-on role for a full-stack developer to create and deliver a range of tools and applications focused on design and deployment of large-scale AI/ML clustered infrastructure. You will be working with the latest agentic tools and patterns to develop, deploy, and maintain these applications. You’ll join a growing team of multi-disciplined engineers that operates across industry verticals as subject matter experts in the AI stack and across the cluster. 

 

THE PERSON: 

  • Demonstrated use of AI coding assistants and LLM-powered developer tools: daily user of AI agents and tools
  • Professional software development experience, including substantial experience building and supporting web applications
  • Proficiency in modern frontend development using JavaScript or TypeScript and a framework such as React, Angular, or Vue
  • Experience developing backend services and APIs using a modern server-side language or framework
  • Experience deploying and operating applications using Linux, containers, CI/CD, and cloud or on-premises infrastructure
  • Experience with automated testing, source control, code review, debugging, and production support
  • Working knowledge of web application security, authentication, authorization, and secure secrets handling  

 

KEY RESPONSIBILITIES: 

  • Partner with engineering peers, domain experts in adjacent teams, and business stakeholders to understand requirements and translate them into flexible, future-proof design solutions
  • Hands on development, iteration, and maintenance of tools and applications that codify various aspects of large-scale AI cluster design stages and cluster deployment activities
  • Design and development of cohesive interface code between disparate third party tools
  • Own features from requirements and design through deployment and ongoing maintenance
  • Work in an iterative software environment, including planning and delivering work in small increments, collaborating with stakeholders, often in different areas of domain expertise (Agile development practices)
  • Participate in code reviews and retros; adapt to changing requirements and priorities

 

PREFERRED EXPERIENCE: 

  • Strong Linux fundamentals: Linux operating systems, networking, filesystems, containers, performance tooling (perf, flamegraphs, nvprof/rocprof, basic eBPF).
  • Clear communication: ability to turn complex systems into accessible, structured documentation with diagrams and reproducible steps
  • AMD ecosystem experience: ROCm, RCCL, Instinct GPUs, EPYC platforms, compiler/toolchain impacts, and performance tuning
  • Orchestration models: Slurm configuration patterns, Kubernetes for HPC/AI (GPU operators, device plugins), Apptainer/Singularity
  • Automation, IaC , and scripting tools/languages (Ansible, Terraform, Python, bash)
  • Storage/data: knowledge of or familiarity with parallel filesystems (Lustre, BeeGFS), object stores, RDMA, data pipeline throughput and caching strategies
  • Hands-on familiarity with on-premises infrastructure, particularly for AI/ML/HPC workloads would be beneficial 

 

ACADEMIC CREDENTIALS: 

  • Bachelors or Masters degree in computer science or software/computer engineering 

 

LOCATIONS:

  • Santa Clara, CA
  • Austin, TX
  • Secaucus, NJ
  • Seattle, WA 

 

This role is not eligible for visa sponsorship.

 

#LI-CB1

#LI-Hybrid



Benefits offered are described:  AMD benefits at a glance.

 

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

 

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD’s “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.



Tags: No, USD $143,500.00/Yr., USD $205,000.00/Yr., US Careers (External)

Get Software Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
About You logo

Frontend Engineer

Hamburg, HH
✓ From careers page· 9m ago
Ledgebrook logo

Head of Forward Deployed Engineering (Remote)

$200k–$250kRemote · US-eligible
✓ From careers page· 9m ago
Sequoia Connect logo

Senior Azure Developer

Remote · Mexico-eligible
✓ From careers page· 1h ago
Sequoia Connect logo

Senior Software Engineer, AI-Assisted

Mexico
✓ From careers page· 1h ago

Frequently asked questions

What is the salary for System Design Engineer, AI Cluster Software Engineer at AMD?

The estimated salary range for System Design Engineer, AI Cluster Software Engineer at AMD is $144,000 - $205,000 USD per year.

What skills are required for System Design Engineer, AI Cluster Software Engineer at AMD?

The required skills for System Design Engineer, AI Cluster Software Engineer at AMD include: AI, Machine Learning, JavaScript, TypeScript, React, Angular, Vue.js, API, Linux, CI/CD, Python, Bash, Agile, Kubernetes, Terraform, Ansible.

What is the seniority level for System Design Engineer, AI Cluster Software Engineer at AMD?

System Design Engineer, AI Cluster Software Engineer at AMD is a Mid Level level position.

How do I apply for System Design Engineer, AI Cluster Software Engineer at AMD?

You can view the full description and apply for System Design Engineer, AI Cluster Software Engineer at AMD on EchoJobs: https://echojobs.io/job/amd-system-design-engineer-ai-cluster-software-engineer-23umk.