1

Hpc System Engineer Jobs (NOW HIRING)

... system configuration for Ethernet, InfiniBand, and Fiber Channel SAN. • Experience with HPC ... programming, desktop support, computer operations, and facilities as required to complete ...

HPC Systems Architect

Chicago, IL · On-site

$200K - $225K/yr

Bachelor's, Master's, or PhD degree in Computer Science, Electrical Engineering, or a related field * Extensive experience (typically 10+ years) in HPC architecture, system design, or a similar role ...

Required : • Expertise in parallel programming models (e.g., MPI, OpenMP) • Strong knowledge of HPC system architectures and workflows • Experience with distributed systems and data-intensive ...

HPC Software Engineer III - UPDATED

Socorro, NM · On-site +1

$54.25 - $72.75/hr

This includes building modern, scalable data-processing systems in partnership with leading high ... Work with HPC system engineers to tune application performance for specific architectures.

Required : • Expertise in parallel programming models (e.g., MPI, OpenMP) • Strong knowledge of HPC system architectures and workflows • Experience with distributed systems and data-intensive ...

HPC Software Engineer III - UPDATED

Green Bank, WV · On-site +1

$46.75 - $63/hr

This includes building modern, scalable data-processing systems in partnership with leading high ... Work with HPC system engineers to tune application performance for specific architectures.

Zealogics Inc is seeking an HPC Engineer ... The primary responsibilities include maintaining and configuring HPC systems, troubleshooting ...

Showing results 41-60

Hpc System Engineer information

See salary details

$53.5K

$127.2K

$167K

How much do hpc system engineer jobs pay per year?

As of Sep 10, 2026, the average yearly pay for hpc system engineer in the United States is $127,215.00, according to ZipRecruiter salary data. Most workers in this role earn between $98,000.00 and $157,000.00 per year, depending on experience, location, and employer.

What is an HPC system engineer?

An HPC (High-Performance Computing) System Engineer designs, deploys, and manages supercomputing environments used for complex computations. They optimize hardware and software components, ensuring system performance, scalability, and reliability. Responsibilities include configuring clusters, troubleshooting performance issues, and maintaining parallel file systems. They work with researchers and developers to optimize code for maximum efficiency. Strong knowledge of Linux, networking, and parallel computing is essential for this role.

What are the key skills and qualifications needed to thrive as an HPC system engineer?

Excelling as an HPC System Engineer requires strong expertise in Linux systems administration, parallel computing, and networking, often supported by a degree in computer science or a related field. Familiarity with HPC resource managers (such as Slurm or PBS), file systems like Lustre or GPFS, and certifications like CompTIA Linux+ or RHCE are highly valuable. Effective problem-solving, teamwork, and communication skills help engineers address complex technical issues and interact with diverse research and engineering teams. These competencies are essential to ensure optimized system performance and support for high-demand computational workloads.

What are some common challenges faced by HPC system engineers?

HPC System Engineers often encounter challenges related to managing large-scale clusters, troubleshooting performance bottlenecks, and ensuring system reliability under demanding workloads. Keeping up with evolving hardware, software updates, and security requirements is also a key part of the job. The role frequently involves responding to urgent issues, supporting a variety of users with different computational needs, and balancing maintenance with ongoing project deadlines. Successfully navigating these challenges requires both strong technical troubleshooting skills and the ability to communicate solutions effectively with researchers and IT peers.

More about Hpc System Engineer jobs

What are the most commonly searched types of Hpc System Engineer jobs?

The most popular types of Hpc System Engineer jobs are:

What are popular job titles related to Hpc System Engineer jobs?

For Hpc System Engineer jobs, the most frequently searched job titles are:

Infographic showing various Hpc System Engineer job openings in the United States as of September 2026, with employment types broken down into 1% As Needed, 85% Full Time, 12% Part Time, and 2% Contract. Highlights an 91% Physical, 2% Hybrid, and 7% Remote job distribution, with an average salary of $127,215 per year, or $61.2 per hour.

HPC Systems Engineer

Houston, TX • On-site

US Tech Solutions
IT Services • 1 - 5K employees

Full-time

Re-posted 11 days ago


Job description

Job Summary:
US Tech Solutions is a global staff augmentation firm providing a wide range of talent on-demand. They are seeking an HPC Systems Engineer to work in a large HPC enterprise environment, responsible for the installation, configuration, and management of Linux-based operating systems and various HPC storage solutions.
Responsibilities:
• A minimum of 5 years’ experience working in a large HPC enterprise environment comprising thousands of servers, large storage solutions, tape and tape automation.
• Proficient in the installation, configuration and management of Linux based operating systems, preferably using RHEL, CentOS, Rocky Linux.
• Experience with IBM’s xCAT distributed computing management software.
• Experience with installation and maintenance of computer hardware including servers, tape drives, robotic tape libraries, GPGPU, SSD, disk arrays.
• Experience with containerization.
• Knowledge of networking and datacenter technologies, switching, routing, high-availability, LAN / WAN / WLAN topologies and system configuration for Ethernet, InfiniBand, and Fiber Channel SAN.
• Experience with HPC Storage Solutions, for example configuration and operation of HPE ClusterStor systems, NetApp, Dell Isilon, and Pure Storage.
• Ability to write and troubleshoot Bourne, Bash and C Shell, Perl, Python, Ruby and MRTG scripts.
• Experience with PostgreSQL and database installation and support.
• Experience with Google Cloud Platform and Azure public clouds. Able to provision and manage instances, build images, write installation scripts.
• Experience with configuration tools like Ansible and Terraform.
• Experience with backup and recovery tools, IBM Spectrum, Dell Networker.
• Good knowledge of Linux security, including configuration of endpoint security tools.
• Ability to evaluate HPC system environments and make recommendations for improvement in performance and manageability.
• Ability to investigate, debug and diagnose system level issues.
• Conform to local change management philosophies, including full testing on non-production systems, prior to production deployment.
• Effectively communicate all change activities to all affected parties including a clear description of the change, related service outages and possible effects on the different environments we support.
• Ensure IT deployment standards are maintained, with verification through reporting systems.
• Meet KPO requirements for InTouch support processing, including full documentation of problem resolution, creation of knowledge content and best practice items.
• Show a good understanding of computer equipment, and its care and maintenance.
• Work with other internal support groups, systems, networking, programming, desktop support, computer operations, and facilities as required to complete administration functions.
• Work with a variety of vendors in technical environments and in the reporting and investigation of system problems.
• Provide a written weekly status report to the team manager and be prepared to present and discuss this with the team at a weekly status meeting.
• Prepared to work outside of normal hours as system maintenance often must be performed outside of prime time; provide 24/7 support to computer operations; work with other remote support locations, for example Kuala Lumpur, backing follow the sun support.
• Participate in support on-call schedule and in weekend power outages, normally two per year and in emergency data center activities.
• Peer-review all major projects, as part of the normal deployment philosophy.
• Ensure compliance with all quality assurance, best practice procedures and QHSE requirements, as defined by job position.
• Self-motivated, able to work with minimum direction.
• Able to work as part of a team, either in small groups, or as part of the Data Center support team as a whole and accomplish this in either a lead or reporting role.
• Able to demonstrate good written, phone and face to face communication skills when working with a peer group and with internal and external customers and with vendors.
• Adhere to industry standard systems administration techniques and procedures.
• Document standard user and operational requirements.
• Willingness to train others.
Qualifications:
Required:
• A minimum of 5 years’ experience working in a large HPC enterprise environment comprising thousands of servers, large storage solutions, tape and tape automation.
• Proficient in the installation, configuration and management of Linux based operating systems, preferably using RHEL, CentOS, Rocky Linux.
• Experience with IBM’s xCAT distributed computing management software.
• Experience with installation and maintenance of computer hardware including servers, tape drives, robotic tape libraries, GPGPU, SSD, disk arrays.
• Experience with containerization.
• Knowledge of networking and datacenter technologies, switching, routing, high-availability, LAN / WAN / WLAN topologies and system configuration for Ethernet, InfiniBand, and Fiber Channel SAN.
• Experience with HPC Storage Solutions, for example configuration and operation of HPE ClusterStor systems, NetApp, Dell Isilon, and Pure Storage.
• Ability to write and troubleshoot Bourne, Bash and C Shell, Perl, Python, Ruby and MRTG scripts.
• Experience with PostgreSQL and database installation and support.
• Experience with Google Cloud Platform and Azure public clouds. Able to provision and manage instances, build images, write installation scripts.
• Experience with configuration tools like Ansible and Terraform.
• Experience with backup and recovery tools, IBM Spectrum, Dell Networker.
• Good knowledge of Linux security, including configuration of endpoint security tools.
• Ability to evaluate HPC system environments and make recommendations for improvement in performance and manageability.
• Ability to investigate, debug and diagnose system level issues.
• Conform to local change management philosophies, including full testing on non-production systems, prior to production deployment.
• Effectively communicate all change activities to all affected parties including a clear description of the change, related service outages and possible effects on the different environments we support.
• Ensure IT deployment standards are maintained, with verification through reporting systems.
• Meet KPO requirements for InTouch support processing, including full documentation of problem resolution, creation of knowledge content and best practice items.
• Show a good understanding of computer equipment, and its care and maintenance.
• Work with other internal support groups, systems, networking, programming, desktop support, computer operations, and facilities as required to complete administration functions.
• Work with a variety of vendors in technical environments and in the reporting and investigation of system problems.
• Provide a written weekly status report to the team manager and be prepared to present and discuss this with the team at a weekly status meeting.
• Prepared to work outside of normal hours as system maintenance often must be performed outside of prime time; provide 24/7 support to computer operations; work with other remote support locations, for example Kuala Lumpur, backing follow the sun support.
• Participate in support on-call schedule and in weekend power outages, normally two per year and in emergency data center activities.
• Peer-review all major projects, as part of the normal deployment philosophy.
• Ensure compliance with all quality assurance, best practice procedures and QHSE requirements, as defined by job position.
• Self-motivated, able to work with minimum direction.
• Able to work as part of a team, either in small groups, or as part of the Data Center support team as a whole and accomplish this in either a lead or reporting role.
• Able to demonstrate good written, phone and face to face communication skills when working with a peer group and with internal and external customers and with vendors.
• Adhere to industry standard systems administration techniques and procedures.
• Document standard user and operational requirements.
• Willingness to train others.
• Strong understanding of Microsoft operating systems
• Excellent PC troubleshooting skills
• Able to read and understand technical manuals
• Ability to multi-task and prioritize projects effectively
• Experience with LAN/WAN networks
• Thorough knowledge of computer systems and IT components.
• Bachelor's degree from a four-year college in computer science studies or 5-10 years equivalent work experience, and current industry recognized training and certification, for example from Cisco, RedHat or Microsoft.
Company:
US Tech Solutions counted among the largest yet the fastest growing staffing firm; all achieved organically. Founded in 2000, the company is headquartered in Toronto, CAN, with a team of 1001-5000 employees. The company is currently Late Stage.

US Tech Solutions logo

About US Tech Solutions

Sourced by ZipRecruiter

US Tech Solutions is a global staff augmentation firm providing a wide range of talent on-demand and total workforce solutions.

Industry

It services

Company size

1,001 - 5,000 Employees

Headquarters location

Jersey City, NJ, US

Year founded

2000

Social media