INflow Federal logo

HPC Infrastructure & Cluster Engineer

INflow Federal

On-site
Springfield, VA
Full-time
Senior
5+ yrs
Salary not listedPosted 59m ago

Real job — pulled straight from INflow Federal’s careers page · Verified September 2, 2026 · No reposts.

Job description

INflow Federal is hiring a HPC Infrastructure & Cluster Engineer — a full-time, based in Springfield, VA role. Apply directly on INflow Federal's careers page below.

HPC Infrastructure & Cluster Engineer

Team: Enterprise Services

Location: Springfield, VA

Commitment: Full-Time Employee

Workplace Type: onsite

Salary:

This is the likely salary range for this position. This is not, however, a guarantee of compensation or salary. Rather, salary will be set based on experience, geographic location and possibly contractual requirements and could fall outside of this range.

About INflow Federal - founded in 2013, INflow Federal is a mission-driven small business delivering cutting-edge solutions to the Department of War (DoW) and Joint Force operations across 20+ states. Our strength comes from our people - especially the Veterans who make up over 50% of our workforce. Through our Veteran Outreach Program and employee-first culture, we invest deeply in professional growth, well-being, and innovation. Known for our agility, transparency, and integrity, INflow combines real-world experience with emerging technologies like AI/ML to help our customers lead in a rapidly evolving defense landscape. We empower both our employees and mission partners to stay ahead - driving smarter, faster, and more secure outcomes.

Job Overview:


We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environmen. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Here, your work is more than a job- it's a journey in innovation. With opportunities to work on high-impact projects, access to the latest technologies, and a culture that thrives on creativity and collaboration, INflow Federal is where your expertise can truly make a difference.

Specific Duties and Responsibilities:

  • Cluster Administration: Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management: Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization: Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management: Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an InfiniBand GPU-to-GPU network infrastructure to minimize latency for distributed operations. 
  • Environment Configuration: Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat OpenShift, required for seamless customer model deployment.
  • Security and Compliance: Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
  •  

Required Skills:

  • Experience: 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Technical Skills:
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with InfiniBand).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
  • Hands-on experience with enterprise container orchestration platforms, specifically OpenShift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus: Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.
  •  

Preferred Skills:

    • Familiarity with parallel file systems and high-throughput storage architectures.
    • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.
Certification Requirements
Most of our contracts require DoD 8140 (formerly 8570) compliance. We built 8140.study - our certification training platform - to help. This helps you study for certifications such as the CompTIA Security+.

Other Notes
  • Some travel may be required: Must have valid driver’s license and transportation. This is subject to change at the direction of the customer.
  • If accommodation is needed with your application or the interview process for applicants with disabilities, please contact Human Resources at 703-594-8601.
  • Candidate must have the ability to lift up to 50 lbs.
  • Must have willingness to perform duties not listed in the job description as required by INflow and our customer.
Citizenship Requirements
* Please note that INflow Federal is a defense contractor. Pursuant to our government contracts, candidates must be US Citizens to be considered for employment.

Equal Opportunity Employer 
Diversity and Inclusion
INflow provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws.
 
This commitment applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, leaves of absence, compensation, and training. Job applicants and employees are evaluated solely on job-related qualifications and experience.
 
 

Get HPC Infrastructure & Cluster Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Weekday logo

Database Administrator

Chennai, TN
✓ From careers page· 35m ago

Frequently asked questions

What skills are required for HPC Infrastructure & Cluster Engineer at INflow Federal?

The required skills for HPC Infrastructure & Cluster Engineer at INflow Federal include: Linux, Bash, Python, OpenShift, Kubernetes, CompTIA Security+.

What is the seniority level for HPC Infrastructure & Cluster Engineer at INflow Federal?

HPC Infrastructure & Cluster Engineer at INflow Federal is a Senior level position.

How do I apply for HPC Infrastructure & Cluster Engineer at INflow Federal?

You can view the full description and apply for HPC Infrastructure & Cluster Engineer at INflow Federal on EchoJobs: https://echojobs.io/job/inflow-federal-hpc-infrastructure-cluster-engineer-uck10.