Jobgether logo

HPC Storage Engineer

Jobgether

Remote
Full-time
Senior
Staff
8+ yrs
$180k–$260kPosted 5d ago

Real job — pulled straight from Jobgether’s careers page · Verified September 12, 2026 · No reposts.

Job description

Jobgether is hiring a HPC Storage Engineer — a full-time, remote role ($180k–$260k). Apply directly on Jobgether's careers page below.

HPC Storage Engineer - West Coast

Team: IT

Location: US

Commitment: Full-time

Workplace Type: remote

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a HPC Storage Engineer - West Coast based in the United States.

This is a senior, hands-on infrastructure engineering role focused on building and scaling a multi-region storage platform for demanding AI workloads.
You’ll own critical storage systems spanning network volumes, local NVMe, and S3-compatible object storage at petabyte scale.
Your work will directly influence training, fine-tuning, and inference performance, including cold-start speed, data streaming, and workload reliability.
You’ll operate at the intersection of storage, networking, hardware, and software, with substantial ownership from architecture through production operations.
The role offers significant latitude to automate manual processes, establish SLOs, optimize performance, and shape long-term storage strategy.
You’ll collaborate closely with SRE, network engineering, supply chain teams, and infrastructure partners in a fast-moving remote environment.
This opportunity is ideal for an engineer who enjoys solving complex production problems and building infrastructure that serves millions of developers.

Accountabilities

    • Own the capacity, durability, availability, and performance of network volumes, local NVMe, and S3-compatible object storage.
    • Tune the complete I/O path, including device and filesystem configuration, caching, read-ahead strategies, replication, erasure coding, and client-side mount behavior.
    • Diagnose complex storage and performance issues end to end, identifying root causes and implementing durable solutions.
    • Lead capacity expansions, hardware refreshes, migrations, and data rebalancing while minimizing or eliminating customer-visible disruption.
    • Design and optimize the networking infrastructure supporting storage workloads, including high-throughput east-west fabrics, MTU and jumbo-frame configuration, congestion and flow control, multipath, and NIC/offload settings.
    • Optimize storage traffic across RDMA/RoCE and high-speed InfiniBand or Ethernet environments, collaborating with network engineering on topology, oversubscription, and cross-region data movement.
    • Develop and ship production software in Go, Python, Rust, or similar languages for storage control-plane services, provisioning, data movement, and monitoring.
    • Build and extend integrations with internal control-plane services, S3-compatible interfaces, CSI drivers, Kubernetes APIs, vendor platforms, and cloud-provider APIs.
    • Replace manual operational procedures with reliable automation and infrastructure-as-code, while participating fully in code reviews, testing, and CI.
    • Instrument storage infrastructure with meaningful metrics covering IOPS, throughput, latency, errors, retries, capacity utilization, and tenant consumption.
    • Build dashboards, SLOs, and alerts that identify degradation proactively and support reliable production operations.
    • Participate in an on-call rotation and lead blameless post-incident follow-through, ensuring lessons learned translate into measurable system improvements.
    • Requirements

      • 8+ years of experience in infrastructure, storage, or systems engineering, including substantial ownership of production storage environments at scale.
      • Deep practical experience with at least one distributed storage platform such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, ZFS-based systems, or a comparable technology.
      • Strong knowledge of Linux internals and the storage stack, including block devices, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI, and NVMe-oF.
      • Hands-on experience building or operating S3-compatible object storage services.
      • Strong networking fundamentals and demonstrated experience tuning networks specifically for storage workloads.
      • Proven ability to write and ship production-quality software using Go, Python, Rust, or a similar programming language, beyond scripting alone.
      • Experience with observability platforms such as Prometheus, Grafana, Datadog, or equivalent, including designing meaningful metrics and monitoring strategies.
      • Demonstrated ability to analyze and resolve performance problems under real production pressure.
      • Self-directed approach, with the ability to take broad infrastructure goals, develop an options analysis, diagnose problems, and execute solutions with minimal supervision.
      • Strong continuous-improvement mindset, with a track record of eliminating operational toil and replacing recurring manual work with automation.
      • High ownership and accountability, including the willingness to follow problems across team boundaries through to resolution.
      • Collaborative, low-ego communication style combined with confidence in technical decision-making.
      • Experience with AI/ML storage workloads, including checkpointing, dataset streaming, model-weight distribution, GPU-adjacent data locality, or GPUDirect Storage, is preferred.
      • Familiarity with Kubernetes storage internals, including CSI drivers, PV/PVC lifecycles, StatefulSets, and local persistent volumes, is a plus.
      • Bare-metal or colocation experience, including hardware selection, vendor management, firmware, and physical failure domains, is beneficial.
      • Experience operating multi-tenant environments where isolation, fairness, and quality of service are critical is preferred.
      • Background in a rapidly scaling cloud or infrastructure provider is advantageous.
      • Benefits

        • Base salary: $180,000–$260,000, with the final range determined based on career level, experience, qualifications, and location.
        • Meaningful equity through stock options, giving employees an opportunity to share in the company’s growth.
        • Generous medical, dental, and vision coverage.
        • Flexible paid time off.
        • Remote-first work environment with collaborative teams and Slack as a primary internal communication channel.
        • $1,200 home office and equipment stipend to help create an effective remote workspace.
        • Opportunity to work on cutting-edge AI infrastructure with a strong emphasis on ownership, learning, and technical impact.
        • Inclusive workplace committed to equal opportunity and respect for people from diverse backgrounds.
        • Candidates must be legally authorized to work in the United States; employment visa sponsorship is not currently available.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
 Why Apply Through Jobgether? 
 
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
 
 
#LI-CL1

Get HPC Storage Engineer jobs like this

New roles from thousands of companies land hourly, straight from their careers pages. Get the freshest matches by email so you never miss one.

Email me new jobs
Clearwater Analytics logo

Site Reliability Engineer

$95k–$123kChicago, IL
✓ From careers page· 41m ago
Endeavour Group logo

Senior Cloud Engineer (Azure)

Sydney, NSW
✓ From careers page· 46m ago
FLPL logo

FLPL

New

IoT Engineer

Mumbai, India
✓ From careers page· 1h ago
Axelera logo

AI Systems Engineer, Agents & Inference

Amsterdam, NL
✓ From careers page· 1h ago

Frequently asked questions

What is the salary for HPC Storage Engineer at Jobgether?

The estimated salary range for HPC Storage Engineer at Jobgether is $180,000 - $260,000 USD per year.

Is HPC Storage Engineer at Jobgether a remote job?

Yes, HPC Storage Engineer at Jobgether is a remote position. This role is open to remote candidates.

What skills are required for HPC Storage Engineer at Jobgether?

The required skills for HPC Storage Engineer at Jobgether include: Linux, S3, Go, Python, Rust, Prometheus, Grafana, Datadog, Kubernetes, Ethernet, AI, Machine Learning.

What is the seniority level for HPC Storage Engineer at Jobgether?

HPC Storage Engineer at Jobgether is a Senior / Staff level position.

How do I apply for HPC Storage Engineer at Jobgether?

You can view the full description and apply for HPC Storage Engineer at Jobgether on EchoJobs: https://echojobs.io/job/jobgether-hpc-storage-engineer-west-coast-0wdp6.