AI/HPC System Engineer
San Jose, CA ยท On-site
AI/HPC System Engineer Location: San Jose, CA (Onsite) Description We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development ...
San Jose, CA ยท On-site
AI/HPC System Engineer Location: San Jose, CA (Onsite) Description We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development ...
San Jose, CA ยท On-site
AI/HPC System Engineer Location: San Jose, CA (Onsite) Description We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development ...
San Jose, CA ยท On-site
$80 - $90/hr
AI/HPC System Engineer Position Description : Protingent Staffing has an exciting contract AI/HPC System Engineer with our client located in San Jose, CA. * We are hiring an AI/HPC System Engineer to ...
San Jose, CA ยท On-site
$80 - $90/hr
AI/HPC System Engineer Position Description : Protingent Staffing has an exciting contract AI/HPC System Engineer with our client located in San Jose, CA. * We are hiring an AI/HPC System Engineer to ...
San Jose, CA ยท On-site
AI/HPC System Engineer Office Location: San Jose, CA Job Type: Full-Time Work Model: Onsite Job Summary: * This role will be responsible for assuring SK hynix emerging technology leadership to be ...
San Jose, CA ยท On-site
AI/HPC System Engineer Office Location: San Jose, CA Job Type: Full-Time Work Model: Onsite Job Summary: * This role will be responsible for assuring SK hynix emerging technology leadership to be ...
San Jose, CA ยท On-site
$175K - $230K/yr
AI/HPC System Engineer Office Location: San Jose, CA Job Type: Full-Time Work Model: Onsite Job Summary: * This role will be responsible for assuring SK hynix emerging technology leadership to be ...
San Jose, CA ยท On-site
$175K - $230K/yr
AI/HPC System Engineer Office Location: San Jose, CA Job Type: Full-Time Work Model: Onsite Job Summary: * This role will be responsible for assuring SK hynix emerging technology leadership to be ...
San Jose, CA ยท On-site
The AI/HPC System Engineer will design, develop, and optimize next-generation memory solutions and system software for energy-efficient AI computing platforms, collaborating closely with hardware and ...
San Jose, CA ยท On-site
The AI/HPC System Engineer will design, develop, and optimize next-generation memory solutions and system software for energy-efficient AI computing platforms, collaborating closely with hardware and ...
Troy, MI ยท On-site
Transforming the Future with the Convergence of Simulation and Data Systems Engineer - HPC Do you like a challenge, are you a complex thinker who likes to solve problems? If so, then you might be the ...
Troy, MI ยท On-site
Transforming the Future with the Convergence of Simulation and Data Systems Engineer - HPC Do you like a challenge, are you a complex thinker who likes to solve problems? If so, then you might be the ...
Troy, MI ยท On-site
Transforming the Future with the Convergence of Simulation and Data Systems Engineer - HPC Do you like a challenge, are you a complex thinker who likes to solve problems? If so, then you might be the ...
Troy, MI ยท On-site
Transforming the Future with the Convergence of Simulation and Data Systems Engineer - HPC Do you like a challenge, are you a complex thinker who likes to solve problems? If so, then you might be the ...
As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale ...
As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale ...
Chicago, IL ยท On-site
Radix Trading is seeking a seasoned HPC Systems Engineer with a passion for Linux, HPC systems and ... Qualifications: * 5+ years of HPC system administration/architecture including RHEL/CentOS/Rocky ...
Chicago, IL ยท On-site
Radix Trading is seeking a seasoned HPC Systems Engineer with a passion for Linux, HPC systems and ... Qualifications: * 5+ years of HPC system administration/architecture including RHEL/CentOS/Rocky ...
Chicago, IL ยท On-site
Radix Trading is seeking a seasoned HPC Systems Engineer with a passion for Linux, HPC systems and ... Qualifications: * 5+ years of HPC system administration/architecture including RHEL/CentOS/Rocky ...
Chicago, IL ยท On-site
Radix Trading is seeking a seasoned HPC Systems Engineer with a passion for Linux, HPC systems and ... Qualifications: * 5+ years of HPC system administration/architecture including RHEL/CentOS/Rocky ...
Annapolis, MD ยท On-site
$150 - $200/hr
Develop test and configuration management strategies tailored for each system. * Contribute to and/or deliver systems engineering documentation at program and system levels. * Propose and evolve HPC ...
Annapolis, MD ยท On-site
$150 - $200/hr
Develop test and configuration management strategies tailored for each system. * Contribute to and/or deliver systems engineering documentation at program and system levels. * Propose and evolve HPC ...
Albuquerque, NM ยท On-site
$125 - $150/hr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
Albuquerque, NM ยท On-site
$125 - $150/hr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
Irving, TX ยท On-site
$125 - $150/hr
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical field, or three (3) years of directly related Linux systems administration or HPC infrastructure ...
Irving, TX ยท On-site
$125 - $150/hr
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical field, or three (3) years of directly related Linux systems administration or HPC infrastructure ...
Menlo Park, CA ยท On-site
$154K - $217K/yr
As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale ...
Menlo Park, CA ยท On-site
$154K - $217K/yr
As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale ...
San Jose, CA ยท On-site
$284K/yr
As an AI/HPC System Architect, your primary responsibility will be to conceptualize the ... D in Electrical and Computer Engineering, or related field with 15+ years of experience in system ...
San Jose, CA ยท On-site
$284K/yr
As an AI/HPC System Architect, your primary responsibility will be to conceptualize the ... D in Electrical and Computer Engineering, or related field with 15+ years of experience in system ...
Albuquerque, NM ยท On-site
$125 - $150/hr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
Albuquerque, NM ยท On-site
$125 - $150/hr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
Irving, TX ยท On-site
Deploying and configuring HPC software (operating systems, job schedulers, parallel programming environments) * Setting up network infrastructure for high-speed data transfer * Performance ...
Irving, TX ยท On-site
Deploying and configuring HPC software (operating systems, job schedulers, parallel programming environments) * Setting up network infrastructure for high-speed data transfer * Performance ...
New York, NY ยท Remote
Troubleshoot application-related issues and ensure minimal disruption to engineering activities ... Perform system health checks, monitoring, and incident tracking for HPC and CAE environments.
Quick apply
New York, NY ยท Remote
Troubleshoot application-related issues and ensure minimal disruption to engineering activities ... Perform system health checks, monitoring, and incident tracking for HPC and CAE environments.
Irving, TX ยท On-site
Deploying and configuring HPC software (operating systems, job schedulers, parallel programming environments) * Setting up network infrastructure for high-speed data transfer * Performance ...
Irving, TX ยท On-site
Deploying and configuring HPC software (operating systems, job schedulers, parallel programming environments) * Setting up network infrastructure for high-speed data transfer * Performance ...
$77K - $105K/yr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
$77K - $105K/yr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
$53.5K - $63.8K
2% of jobs
$63.8K - $74.1K
4% of jobs
$74.1K - $84.5K
7% of jobs
$84.5K - $94.8K
9% of jobs
$97.6K is the 25th percentile. Wages below this are outliers.
$94.8K - $105.1K
10% of jobs
$105.1K - $115.4K
7% of jobs
$115.4K - $125.7K
10% of jobs
The median wage is $127.4K / yr.
$125.7K - $136K
6% of jobs
$136K - $146.4K
3% of jobs
$156.4K is the 75th percentile. Wages above this are outliers.
$146.4K - $156.7K
17% of jobs
$156.7K - $167K
24% of jobs
$53.5K
$127.2K
$167K
An HPC (High-Performance Computing) System Engineer designs, deploys, and manages supercomputing environments used for complex computations. They optimize hardware and software components, ensuring system performance, scalability, and reliability. Responsibilities include configuring clusters, troubleshooting performance issues, and maintaining parallel file systems. They work with researchers and developers to optimize code for maximum efficiency. Strong knowledge of Linux, networking, and parallel computing is essential for this role.
Excelling as an HPC System Engineer requires strong expertise in Linux systems administration, parallel computing, and networking, often supported by a degree in computer science or a related field. Familiarity with HPC resource managers (such as Slurm or PBS), file systems like Lustre or GPFS, and certifications like CompTIA Linux+ or RHCE are highly valuable. Effective problem-solving, teamwork, and communication skills help engineers address complex technical issues and interact with diverse research and engineering teams. These competencies are essential to ensure optimized system performance and support for high-demand computational workloads.
HPC System Engineers often encounter challenges related to managing large-scale clusters, troubleshooting performance bottlenecks, and ensuring system reliability under demanding workloads. Keeping up with evolving hardware, software updates, and security requirements is also a key part of the job. The role frequently involves responding to urgent issues, supporting a variety of users with different computational needs, and balancing maintenance with ongoing project deadlines. Successfully navigating these challenges requires both strong technical troubleshooting skills and the ability to communicate solutions effectively with researchers and IT peers.
The most popular types of Hpc System Engineer jobs are:
For Hpc System Engineer jobs, the most frequently searched job titles are:

San Jose, CA โข On-site
Other
Posted 16 days ago
Position Title: AI/HPC System Engineer
Location: San Jose, CA (Onsite)
Description
We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development workloads. This role deploys, automates, and maintains GPU clusters across on-premise and cloud environments, delivering reliable, scalable, and cost-efficient compute for engineering and R&D teams.
Responsibilities:
โข GPU/HPC infrastructure: Build, configure, and operate GPU and HPC clusters across compute, storage, and networking; support capacity planning, performance tuning, and optimization for AI training, inference, and compute-intensive workloads
โข Hybrid cloud infrastructure: Deploy and maintain compute environments spanning on-premise and public cloud, and contribute to modernization and scaling initiatives for HPC/AI infrastructure
โข Automation and observability: Implement infrastructure-as-code, provisioning automation, monitoring, and alerting, and drive improvements in resource utilization and efficiency
โข AI platform support: Deploy, integrate, and support LLM APIs, coding assistants, and AI/agent platforms used by internal engineering teams
โข Operations and collaboration: Troubleshoot and resolve infrastructure issues, document standards and runbooks, and work with relevant stakeholders to support day-to-day IT operations
Qualifications:
โข Bachelor''s degree in Computer Science, Engineering, or a related technical field
โข 3+ years of hands-on experience in IT infrastructure, cloud, platform engineering, or HPC
โข Hands-on experience with Linux-based infrastructure and public cloud environments such as AWS, Azure, or Google Cloud Platform
โข Experience deploying or operating GPU/HPC environments, including workload scheduling or orchestration platforms such as Kubernetes or Slurm
โข Experience with infrastructure automation, monitoring, troubleshooting, and performance optimization
โข Solid understanding of compute, storage, networking, and container technologies; experience with AI/ML infrastructure or workloads is a plus
โข Strong collaboration and communication skills, with the ability to work across engineering and IT teams
Sourced by ZipRecruiter
It services and software development
1,001 - 5,000 Employees
Sunnyvale, CA, US