Staff AI Infrastructure Engineer
$241K - $331K/yr
Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the ...
$241K - $331K/yr
Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the ...
$241K - $331K/yr
Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the ...
San Francisco, CA · On-site
$126K - $166K/yr
You'll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the ...
San Francisco, CA · On-site
$126K - $166K/yr
You'll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the ...
San Francisco, CA · On-site
$150 - $200/hr
Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads * Build software ...
San Francisco, CA · On-site
$150 - $200/hr
Extend scheduling and orchestration systems such as Kubernetes and Slurm for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads * Build software ...
Livingston, NJ · On-site
$153K - $204K/yr
HPC and fabric verification, Slurm-on-Kubernetes (SUNK), and further down the stack into firmware and hardware, building the abstractions that make new coverage easy to add. * Keep the pipeline fast ...
Livingston, NJ · On-site
$153K - $204K/yr
HPC and fabric verification, Slurm-on-Kubernetes (SUNK), and further down the stack into firmware and hardware, building the abstractions that make new coverage easy to add. * Keep the pipeline fast ...
Murphy, TX · On-site
Experience with Slurm or similar schedulers * Preferred: Exposure to HPCM or parallel file systems
Murphy, TX · On-site
Experience with Slurm or similar schedulers * Preferred: Exposure to HPCM or parallel file systems
Bellevue, WA · On-site
$153K - $204K/yr
HPC and fabric verification, Slurm-on-Kubernetes (SUNK), and further down the stack into firmware and hardware, building the abstractions that make new coverage easy to add. * Keep the pipeline fast ...
Quick apply
Bellevue, WA · On-site
$153K - $204K/yr
HPC and fabric verification, Slurm-on-Kubernetes (SUNK), and further down the stack into firmware and hardware, building the abstractions that make new coverage easy to add. * Keep the pipeline fast ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
Murphy, TX · On-site
Experience with Slurm or similar schedulers * Preferred: Exposure to HPCM or parallel file systems
Murphy, TX · On-site
Experience with Slurm or similar schedulers * Preferred: Exposure to HPCM or parallel file systems
Murphy, TX · On-site
Experience with Slurm or similar schedulers * Preferred: Exposure to HPCM or parallel file systems
Murphy, TX · On-site
Experience with Slurm or similar schedulers * Preferred: Exposure to HPCM or parallel file systems
San Jose, CA · On-site
$120K - $150K/yr
Make Slurm the primary scheduler and operational control plane for interactive, batch, regression, and multi-day simulation workloads. Deploy and operate Slurm directly and/or through AWS ...
San Jose, CA · On-site
$120K - $150K/yr
Make Slurm the primary scheduler and operational control plane for interactive, batch, regression, and multi-day simulation workloads. Deploy and operate Slurm directly and/or through AWS ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
Murphy, TX · On-site
Experience with Slurm or similar schedulers * Preferred: Exposure to HPCM or parallel file systems
Murphy, TX · On-site
Experience with Slurm or similar schedulers * Preferred: Exposure to HPCM or parallel file systems
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex ...
$11K - $29.3K
0% of jobs
$37.5K is the 25th percentile. Wages below this are outliers.
$29.3K - $47.6K
56% of jobs
$47.6K - $66K
0% of jobs
$66K - $84.3K
0% of jobs
$84.3K - $102.6K
0% of jobs
$102.6K - $120.9K
0% of jobs
$120.9K - $139.2K
14% of jobs
$149.5K is the 75th percentile. Wages above this are outliers.
$139.2K - $157.5K
8% of jobs
$157.5K - $175.9K
0% of jobs
$175.9K - $194.2K
0% of jobs
$194.2K - $212.5K
21% of jobs
$11K
$106.4K
$212.5K
Cities with the most Slurm job openings:
States with the most job openings for Slurm jobs include:
The top searched job categories for Slurm jobs are:

$241K - $331K/yr
Full-time
Retirement, PTO
Re-posted 7 days ago
Own reliability, observability, and incident response for multi-site GPU clusters running Slurm on Kubernetes.
Debug and resolve deep infrastructure failures across storage, networking, scheduling, and GPU compute layers.
Design and execute GPU cluster scaling plans and build automation and tooling to manage cluster operations at scale.
The AI Cluster Production Engineering team is part of the AI Compute Platform organization at Biohub, a non-profit research lab committed to open science and open-source AI. We own the design, operation, and reliability of large-scale multi-GPU AI clusters that power frontier AI biology research: protein language models, genomic foundation models, and scientific reasoning systems built to be shared, not monetized. Our clusters run Slurm on Kubernetes infrastructure and support everything from day-to-day AI researcher workflows to multi-node hero training runs at thousands of GPUs. The team works at the intersection of AI tooling, distributed systems, HPC, and frontier AI, debugging deep AI infrastructure problems and building AI systems critical to the entire AI organization.
The OpportunityCZ Biohub's mission is to cure or prevent all human disease. Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with the twist that the mission is biology and the science is open. You'll keep GPU clusters running at high utilization, debug the toughest distributed systems failures, and build the operational foundations for scaling to multi-thousand GPU hero runs. The technical problems are genuinely hard (e.g., multi-node distributed training, InfiniBand fabrics, large-scale storage, Slurm at scale) inside an organization where the work is aimed at helping people, not optimizing ad revenue.
What You'll DoThe Redwood City, CA base pay range for a new hire in this role is $241,000 - $331,000. New hires are typically hired into the lower portion of the range, enabling employee growth in the range over time. Actual placement in range is based on job-related skills and experience, as evaluated throughout the interview process.Â
Better TogetherAs we grow, we're excited to strengthen in-person connections and cultivate a collaborative, team-oriented environment. This role is a hybrid position requiring you to be onsite for at least 60% of the working month, approximately 3 days a week, with specific in-office days determined by the team's manager. The exact schedule will be at the hiring manager's discretion and communicated during the interview process.
Benefits for the Whole YouÂWe're thankful to have an incredible team behind our work. To honor their commitment, we offer a wide range of benefits to support the people who make all we do possible.Â
If you're interested in a role but your previous experience doesn't perfectly align with each qualification in the job description, we still encourage you to apply as you may be the perfect fit for this or another role.
#LI-HybridÂ