1

Hpc Infrastructure Engineer Jobs (NOW HIRING)

HPC/ML Infrastructure Engineer

San Francisco, CA · On-site

$126K - $166K/yr

Spellbrush is seeking an experienced HPC/ML Infrastructure Engineer to lead the administration and operations of a large anime AI training cluster. The role involves bridging the gap between ...

HPC Infrastructure Platform Engineer Founded in 1999 in the beautiful Smoky Mountains of East Tennessee, Cadre5 provides innovative technical solutions to our customers locally and nationally. Our ...

... HPC Infrastructure Platform Engineer to join the HPC Infrastructure group. The preferred candidate will possess commensurate knowledge, skills and abilities in addition to relevant education ...

We are seeking an entry-level (or Junior-level) Infrastructure / HPC Support Engineer to join our growing team. This role focuses on solution deployment, system support, and hardware troubleshooting ...

HPC/ML Infrastructure Engineer

San Francisco, CA · On-site

$126K - $166K/yr

We're looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You'll serve as the ...

HPC/ML Infrastructure Engineer

San Francisco, CA · On-site

$126K - $166K/yr

We're looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You'll serve as the ...

AI & HPC Infrastructure Engineer

Saint Louis, MO · On-site

$97K - $127K/yr

Experience managing the deployment of 1,000+ GPU clusters for AI, HPC, and agentic AI workloads with infrastructure services enabled. * Design and build experience in AI Cloud platforms from ...

Showing results 21-40

Hpc Infrastructure Engineer information

See salary details

$46.5K

$127.1K

$182K

How much do hpc infrastructure engineer jobs pay per year?

As of Sep 9, 2026, the average yearly pay for hpc infrastructure engineer in the United States is $127,066.00, according to ZipRecruiter salary data. Most workers in this role earn between $107,500.00 and $141,000.00 per year, depending on experience, location, and employer.

What is an HPC Infrastructure Engineer?

HPC Infrastructure Engineers are professionals who design, implement, and manage high-performance computing (HPC) systems and environments. They ensure that computational clusters, storage solutions, and networks are optimized for maximum performance and reliability. These engineers often work closely with researchers, scientists, and IT staff to support complex computing workloads, troubleshoot issues, and maintain system security. Their role is crucial in environments where large-scale simulations, data analysis, or scientific research require significant computational power.

What are some common challenges faced by an HPC Infrastructure Engineer when maintaining high-performance computing clusters?

HPC Infrastructure Engineers often encounter challenges related to scalability, hardware failures, and network bottlenecks when maintaining large computing clusters. Keeping systems up-to-date without causing downtime, troubleshooting complex performance issues, and ensuring efficient resource allocation are daily concerns. Collaborating closely with system administrators, researchers, and software engineers is essential to balance user demands with cluster stability and security. Proactive monitoring and clear documentation can help manage these challenges effectively.

What are the key skills and qualifications needed to thrive as an HPC Infrastructure Engineer?

To thrive as an HPC Infrastructure Engineer, you need expertise in computer science, Linux system administration, networking, and parallel computing, often supported by a relevant degree or industry certifications. Familiarity with HPC cluster management tools (like Slurm or PBS), scripting languages (such as Python or Bash), and hardware management is essential. Strong problem-solving abilities, communication skills, and teamwork are valuable soft skills in this role. These skills ensure the efficient deployment, maintenance, and optimization of high-performance computing systems crucial for research and enterprise applications.

What is the difference between Hpc Infrastructure Engineer vs Network Engineer?

AspectHpc Infrastructure EngineerNetwork Engineer
Required CredentialsBachelor's in Computer Science, Engineering, or related field; certifications like Cisco CCNA or CompTIA Network+Bachelor's in Computer Science, Engineering, or related field; certifications like Cisco CCNA or CompTIA Network+
Work EnvironmentData centers, research labs, high-performance computing clustersCorporate offices, data centers, network operation centers
Employer & Industry UsageResearch institutions, tech companies, scientific organizationsTelecommunications, IT service providers, large enterprises

While both roles require networking knowledge and certifications, the Hpc Infrastructure Engineer specializes in managing high-performance computing systems and clusters, whereas the Network Engineer focuses on designing and maintaining general network infrastructure. The Hpc role is more research and data-intensive, often in scientific or academic settings, while Network Engineers work across various industries to ensure connectivity and network security.

Are HPC Infrastructure engineers in demand?

HPC Infrastructure engineers are in high demand due to the growing need for powerful computing systems in research, scientific, and enterprise environments. They require expertise in cluster management, networking, and tools like Linux and job schedulers, making their skills highly sought after in industries relying on high-performance computing. Employment opportunities are expected to grow as data processing and computational needs increase across sectors.

How much do HPC Infrastructure engineers make in the US?

HPC Infrastructure engineers in the US typically earn between $80,000 and $130,000 annually, depending on experience, location, and certifications. Senior roles or those with specialized skills in high-performance computing environments can earn higher salaries, often exceeding $150,000.

What are popular job titles related to Hpc Infrastructure Engineer jobs?

For Hpc Infrastructure Engineer jobs, the most frequently searched job titles are:

Infographic showing various Hpc Infrastructure Engineer job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 92% Full Time, 3% Part Time, and 4% Contract. Highlights an 85% Physical, 4% Hybrid, and 11% Remote job distribution, with an average salary of $127,066 per year, or $61.1 per hour.

HPC/ML Infrastructure Engineer

San Francisco, CA • On-site

Spellbrush
Software Development • 1 - 10 employees

$126K - $166K/yr

Full-time

Re-posted 26 days ago


Job description

Job Summary:
Spellbrush is seeking an experienced HPC/ML Infrastructure Engineer to lead the administration and operations of a large anime AI training cluster. The role involves bridging the gap between researchers and GPU machines, ensuring efficient job running and model training while managing the complex HPC software landscape.
Responsibilities:
• Lead bringup, administration, and operations on the largest anime AI training cluster.
• Serve as the bridge between researchers and GPU machines.
• Ensure SLURM jobs are running, parallel filesystems are serving, and networks are transmitting.
• Manage the modern HPC software landscape including SLURM, K8s, warewulf, MAAS, ansible, WEKA, VAST, Ceph, VPN, and monitoring tools.
• Bring up and manage clusters with traditional Linux sysadmin skills.
• Work on physical computers and assist in building out edge datacenters.
• Collaborate with a small, fast-paced research team.
Qualifications:
Required:
• Experience in HPC infrastructure engineering
• Familiarity with modern HPC software landscape
• Experience with SLURM, Slinky on K8s, warewulf/MAAS/ansible, WEKA/VAST/Ceph
• Experience with VPN and access through tailscale
• Experience with monitoring via Grafana/Prometheus stack
• Linux sysadmin skills including wrangling ldap, triaging dmesg, and setting sticky bits on directories
• Comfortable working with physical computers
• Ability to work on small, fast-paced teams
• Willingness to work on-site in either Tokyo or San Francisco
Preferred:
• Love for anime and the anime aesthetic
• Experience in building and managing edge datacenters
Company:
Spellbrush is an illustration platform that enables artists, illustrators, and animators to complete projects more quickly. Founded in 2016, the company is headquartered in San Francisco, USA, with a team of 11-50 employees. The company is currently Early Stage.