
Real job — pulled straight from Jobgether’s careers page · Verified August 12, 2026 · No reposts.
Job description
Jobgether is hiring a Senior HPC Cluster Engineer, AI/ML — a full-time, based in India role. Apply directly on Jobgether's careers page below.
Senior HPC Cluster Engineer - AI, ML
Team: IT
Location: India
Commitment: Full-time
Workplace Type: remote
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior HPC Cluster Engineer - AI, ML based in India.
This role offers the opportunity to engineer and operate large-scale infrastructure powering advanced AI and high-performance computing workloads. You will help build and manage heterogeneous GPU-accelerated clusters across on-premises and cloud environments. The position combines production operations, automation, performance engineering, reliability, and direct collaboration with AI/ML researchers. You will play a key role in improving resource utilization, system stability, and the overall experience of users running demanding workloads. Working with globally distributed engineering teams, you will influence infrastructure strategy and continuously improve operational practices. The environment is highly technical, fast-moving, and focused on enabling the next generation of AI and scientific computing.
Accountabilities:
- Provide technical leadership for systems administration and service delivery across large-scale AI/HPC infrastructure, coordinating upgrades, incident response, and reliability improvements.
- Own the day-to-day operation of production AI/HPC clusters, monitoring system health, user experience, resource utilization, and adherence to internal service-level targets.
- Build, maintain, and scale heterogeneous AI/ML clusters across on-premises and cloud environments, covering compute, networking, storage, and GPU-accelerated infrastructure.
- Develop scalable automation and tooling to improve the deployment, configuration, management, and operational efficiency of AI/HPC environments.
- Collaborate with global engineering teams and internal users to understand evolving research and computing requirements and deliver reliable infrastructure solutions.
- Support researchers in running complex AI/HPC workloads, including performance analysis, troubleshooting, optimization, and workload tuning.
- Analyze cluster efficiency, job fragmentation, and GPU utilization to identify opportunities to reduce resource waste and improve overall capacity.
- Conduct root cause analysis for infrastructure and service issues, proactively identifying risks and implementing corrective actions before they affect users.
- Lead SEV triage, incident response, and postmortems for reliability events affecting production infrastructure or users.
- Build strong relationships with customers, researchers, and cross-functional engineering teams to improve service delivery and anticipate infrastructure needs.
- Participate in an on-call rotation and provide timely support for critical production GPU clusters.
- Bachelor's degree in Computer Science, Electrical Engineering, or a related discipline, or equivalent practical experience.
- 5+ years of experience designing, deploying, and operating large-scale compute or infrastructure environments.
- Strong experience with AI/HPC job schedulers such as Slurm, Kubernetes, PBS, RTDA, BCM, or LSF.
- Proficiency administering Linux distributions such as CentOS/RHEL and/or Ubuntu.
- Hands-on experience with cluster configuration and infrastructure management tools including BCM, Terraform, Ansible, Puppet, Salt, or similar technologies.
- Strong knowledge of container technologies such as Docker, Singularity, Podman, Shifter, or Charliecloud.
- Proficiency with Python programming and Bash scripting for infrastructure automation and operational tooling.
- Practical experience supporting AI/HPC workflows that use MPI.
- Experience analyzing and tuning the performance of diverse AI and HPC workloads.
- Strong troubleshooting, analytical, and root-cause analysis skills, with the ability to resolve complex distributed infrastructure problems.
- Excellent collaboration and communication skills, particularly when working with researchers, infrastructure engineers, and globally distributed teams.
- Passion for continuous learning and for staying current with emerging technologies and best practices in HPC and AI/ML infrastructure.
- Experience working with NVIDIA GPUs, CUDA programming, NCCL, or MLPerf benchmarking.
- Understanding of AI/ML concepts, algorithms, models, and frameworks such as PyTorch or TensorFlow.
- Experience with InfiniBand, IPoIB, and RDMA technologies.
- Knowledge of high-performance distributed storage platforms such as Lustre or GPFS.
- Experience with advanced GPU cluster optimization and large-scale AI/ML infrastructure.
- Opportunity to work on large-scale AI and high-performance computing infrastructure supporting advanced research and engineering workloads.
- Exposure to GPU-accelerated computing, distributed systems, cloud infrastructure, HPC, and emerging AI/ML technologies.
- Collaboration with highly skilled engineering and research teams across global locations.
- Opportunities to influence infrastructure architecture, automation, reliability, and operational best practices.
- Continuous learning and exposure to rapidly evolving technologies in AI, ML, and high-performance computing.
- Flexible work arrangements may be available, with opportunities based in Bengaluru, Pune, or remote within India.
- Full-time role within a technically advanced and innovation-focused environment.
- Opportunity to contribute to infrastructure that enables next-generation AI and scientific computing.
Requirements:
Nice-to-have experience:
Benefits:
Get Senior HPC Cluster Engineer, AI/ML jobs like this→
New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.
Email me new jobsSimilar jobs




Frequently asked questions
What skills are required for Senior HPC Cluster Engineer, AI/ML at Jobgether?
The required skills for Senior HPC Cluster Engineer, AI/ML at Jobgether include: Kubernetes, Terraform, Ansible, Puppet, Docker, Python, Bash, PyTorch, TensorFlow.
What is the seniority level for Senior HPC Cluster Engineer, AI/ML at Jobgether?
Senior HPC Cluster Engineer, AI/ML at Jobgether is a Senior level position.
How do I apply for Senior HPC Cluster Engineer, AI/ML at Jobgether?
You can view the full description and apply for Senior HPC Cluster Engineer, AI/ML at Jobgether on EchoJobs: https://echojobs.io/job/jobgether-senior-hpc-cluster-engineer-ai-ml-3eftr.