1

Hpc Systems Engineer Jobs in California (NOW HIRING)

AI/HPC Systems Engineer

San Jose, CA · On-site

$150 - $200/hr

AI/HPC Systems Engineer Position Overview We are seeking an AI/HPC Systems Engineer to build, deploy, and operate the compute infrastructure supporting high-performance computing and AI development ...

AI/HPC Systems Engineer Position Overview We are seeking an AI/HPC Systems Engineer to build, deploy, and operate the compute infrastructure supporting high-performance computing and AI development ...

As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex ...

As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex ...

As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex ...

AI/HPC System Engineer

San Jose, CA · On-site

$80 - $90/hr

AI/HPC System Engineer Position Description : Protingent Staffing has an exciting contract AI/HPC System Engineer with our client located in San Jose, CA. * We are hiring an AI/HPC System Engineer to ...

next page

Showing results 1-20

Hpc Systems Engineer information

What is an HPC systems engineer?

HPC Systems Engineers are professionals who design, deploy, manage, and optimize high-performance computing (HPC) systems. These systems are used for complex computational tasks in fields such as scientific research, engineering, and data analytics. HPC Systems Engineers ensure that computing clusters, storage solutions, and networking components run efficiently, securely, and reliably. They also support users, troubleshoot issues, and often work with specialized software, hardware, and parallel processing technologies.

What are the key skills and qualifications needed to thrive as an HPC systems engineer?

To thrive as an HPC Systems Engineer, you need a solid background in computer science, Linux/Unix system administration, and parallel computing concepts, often supported by a relevant degree. Experience with HPC cluster management tools, workload schedulers (like Slurm or PBS), and programming languages such as Python or C/C++ is typically required. Strong problem-solving skills, attention to detail, and the ability to communicate technical concepts clearly are essential soft skills. These qualifications ensure efficient deployment, maintenance, and optimization of high-performance computing infrastructure critical for research and enterprise workloads.

What are some common challenges an HPC systems engineer faces when supporting large-scale computing clusters?

HPC Systems Engineers often encounter challenges such as managing the complexity of cluster architectures, ensuring optimal performance, and troubleshooting hardware or software issues in high-demand environments. They must also coordinate downtime for maintenance without disrupting critical computations and stay updated on rapidly evolving HPC technologies. Strong collaboration with researchers, IT staff, and vendors is essential to address user needs and implement effective solutions.

What is the difference between Hpc Systems Engineer vs Hpc Network Engineer?

AspectHpc Systems EngineerHpc Network Engineer
CredentialsTypically requires a degree in computer science, engineering, or related field; certifications like Cisco CCNA or Linux certifications are commonSimilar credentials; often holds networking certifications such as Cisco CCNP or CompTIA Network+
Work EnvironmentWorks on high-performance computing systems, hardware, and software integration in research or enterprise data centersFocuses on designing, implementing, and maintaining HPC network infrastructure within data centers or research facilities
Industry UsageUsed in scientific research, academia, and enterprise sectors with HPC needsCommon in data centers, research institutions, and organizations requiring advanced network performance

Hpc Systems Engineers and Hpc Network Engineers share overlapping skills in hardware, software, and certifications. However, Hpc Systems Engineers focus on overall system setup and management, while Hpc Network Engineers specialize in network infrastructure. Both roles are vital in supporting high-performance computing environments.

What cities in California are hiring for Hpc Systems Engineer jobs?

Cities in California with the most Hpc Systems Engineer job openings:

Infographic showing various Hpc Systems Engineer job openings in California as of August 2026, with employment types broken down into 80% Full Time, 14% Part Time, 5% Contract, and 1% Nights. Highlights an 86% Physical, 2% Hybrid, and 12% Remote job distribution.

HPC Systems Engineer

San Francisco, CA

University of California San Francisco
Colleges, Universities, and Professional Schools • 10K+ employees

Full-time

Re-posted 5 days ago


University Of California San Francisco rating

7.8

Company rating: 7.8 out of 10

Based on 13 frontline employees who took The Breakroom Quiz

236th of 633 rated colleges and universities


Job description

The CoreHPC team at UCSF is seeking an HPC Systems Engineer to play a key role in the development, maintenance, and day-to-day operations of the Institute's HPC clusters. 

The HPC Systems Engineer will: 

  • Apply advanced systems infrastructure concepts and skills to the operations and improvement of large-scale and highly complex research Cyber Infrastructure (CI) with unique computing, networking, and storage systems designed to address cutting-edge research problems
  • Apply their engineering and design skills to develop new CI solutions, to develop and enhance monitoring to maintain the integrity of CI systems. 
  • Select methods, techniques and evaluation criteria to develop new CI solutions to address complex research problems.
  • Be an active member of the support and maintenance efforts for the CoreHPC cluster, resolving user issues, fixing technical problems, resolving outages, patching, and maintaining systems' uptime and availability. 
  • Provides consultation, support, and guidance to researchers on how to address computational problems using standard tools, packages, and approaches. 
  • Develop enhancements of monitoring to maintain the integrity of CI systems.
  • Participate in multiple technical projects simultaneously. 
  • Applies working knowledge of security control frameworks to maintain the integrity of the CI systems and the research being performed on them. 
  • Gives presentations to the associated team and other technical units. 
  • Evaluates new technologies, including performing moderate to complex cost/benefit analyses.

This position may lead to cross-functional technical working groups and projects in support of onboarding research customers, or making systems improvements. 

Department Overview 

Academic Research Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository of HIPAA-compliant clinical and life sciences data and a centralized, secure, professionally managed infrastructure for the storage and management of research data. ARS empowers medical scientific investigations by offering secure computing environments, data capture, management and analysis tools, and support services which meet researchers' needs. 

The Core HPC team of the Academic Research Service (ARS) focuses on large-scale, high-performance computational and storage services for UCSF researchers so they can address complex computational, AI,  and data science problems.

About UCSF

The University of California, San Francisco (UCSF) is a leading university dedicated to promoting health worldwide through advanced biomedical research, graduate-level education in the life sciences and health professions, and excellence in patient care. It is the only campus in the 10-campus UC system dedicated exclusively to the health sciences. We bring together the world's leading experts in nearly every area of health. We are home to five Nobel laureates who have advanced the understanding of cancer, neurodegenerative diseases, aging and stem cells.

Pride Values

UCSF is a diverse community made of people with many skills and talents. We seek candidates whose work experience or community service has prepared them to contribute to our commitment to professionalism, respect, integrity, diversity and excellence - also known as our PRIDE values.

In addition to our PRIDE values, UCSF is committed to equity - both in how we deliver care as well as our workforce. We are committed to building a broadly diverse community, nurturing a culture that is welcoming and supportive, and engaging diverse ideas for the provision of culturally competent education, discovery, and patient care. Additional information about UCSF is available here.

Join us to find a rewarding career contributing to improving healthcare worldwide.

Equal Employment Opportunity

The University of California is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, age, protected veteran status, or other protected status under state or federal law.

Salary Information

The final salary and offer components are subject to additional approvals based on UC policy.

Your placement within the salary range is dependent on a number of factors including your work experience and internal equity within this position classification at UCSF. For positions that are represented by a labor union, placement within the salary range will be guided by the rules in the collective bargaining agreement.

To learn more about the benefits of working at UCSF, including total compensation, please visit: https://ucnet.universityofcalifornia.edu/compensation-and-benefits/index.html

REQUIRED QUALIFICATIONS

  • Bachelor's degree in a related area such as computer science or engineering, and 6+ years of experience with large-scale or HPC systems * or*  10+ years of related experience with large-scale or HPC systems
  • Expert knowledge of HPC systems infrastructure design
  • Strong knowledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc.
  • Advanced knowledge of computer security best practices and policies including demonstrated experience securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requirements
  • Demonstrated testing and test planning skills. Demonstrated ability to create automated testing.
  • Knowledge of HPC job scheduler system design and operation such as SLURM or PBS, 
  • Demonstrated skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband based clusters 
  • Ability to elicit and communicate technical and non-technical information in a clear and concise manner.
  • Self-motivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to learn effectively and meet deadlines.
  • Understanding of system performance monitoring and actions that can be taken to improve or correct performance.
  • Demonstrated advanced knowledge, skills and abilities associated with system problem identification and resolution. Experience with design, configuration, operation, repair, and tuning of technology systems.
  • Advanced experience writing and editing the most complex scripts used to perform system maintenance and administration.
  • Ability to write technical documentation in a clear and concise manner. Ability to develop runbooks defining complex technical processes in a clear and concise manner 

PREFERRED QUALIFICATIONS

  • Knowledge of the design, development, and application of technology and systems to meet business needs.
  • General knowledge of other areas of IT. Thorough understanding of and experience with systems-related issues and actions that can be taken to improve or correct performance.
  • Demonstrated skills associated with adapting equipment and technology to serve user needs. Demonstrated comprehensive understanding of how system management actions affect other systems, system users and dependent/related functions.

of time

Essential Function (Yes/No)

  

Key Responsibilities

(To be completed by Supervisor)

15

 Applies advanced systems / infrastructure concepts to define, design, implement, and operate highly complex, research cyberinfrastructure systems, services and technology solutions. Proposes and implements highly complex system or device enhancements such as software, hardware and network configuration, updates and installations for projects or services of broad scope. Sets standards for monitoring and maintaining the health and integrity of CI systems including upgrading and patching.

15

 Independently manages systems and services for a large facility, campuswide, medical center or Office of the President and / or institution-wide scope and makes recommendations for purchases or upgrades. Performs complex and advanced analysis to acquire, install, modify and support operating systems, databases, utilities and web-related tools. Selects methods and techniques to obtain solutions. Interacts with senior management. May perform complex network integration tasks and interoperability assessments for interconnected servers or components of clusters for communication. Support and collaborate with researchers and other key IT (e.g. network and security) and Data Center partners in a timely manner

15

 Specifies, writes and executes highly complex software and scripts to support systems management, log analysis, monitoring, deployment, configuration management, and other system administration duties for multiple, highly integrated systems.

30

 Provides consultation, training, support, and guidance to researchers enabling them to utilize  HPC resources effectively. 

10

 Maintains complex security systems. Interprets and adopts campus, medical center or Office of the President, system and regulation-based security policies to control access to networked resources. Provides recommendations and requirements on network access controls.

5

 Collaborates and may provide leadership with other Systems Engineers within the CI ecosystem/higher-education community. Regularly contribute best practices documentation, present at conferences, or publish in peer reviewed journals.

10

 Define and track performance metrics to ensure efficient current and future use of cyber infrastructure resources.

100%

 (To update total %, enter the amount of time in whole numbers (without the % symbol - e.g., 15, 20) then highlight the total sum (e.g., 1%) at the bottom of the column and press F9. The total sum should add up to 100%.)

What University Of California San Francisco employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom