The HPC Systems Engineer will: * Apply advanced systems infrastructure concepts and skills to the operations and improvement of large-scale and highly complex research Cyber Infrastructure (CI) with ...
The HPC Systems Engineer will: * Apply advanced systems infrastructure concepts and skills to the operations and improvement of large-scale and highly complex research Cyber Infrastructure (CI) with ...
HPC Systems Engineer
San Francisco, CA · On-site
The HPC Systems Engineer will: * Apply advanced systems infrastructure concepts and skills to the operations and improvement of large-scale and highly complex research Cyber Infrastructure (CI) with ...
HPC Systems Engineer
San Francisco, CA · On-site
The HPC Systems Engineer will: * Apply advanced systems infrastructure concepts and skills to the operations and improvement of large-scale and highly complex research Cyber Infrastructure (CI) with ...
AI/HPC Systems Engineer
San Jose, CA · On-site
$150 - $200/hr
AI/HPC Systems Engineer Position Overview We are seeking an AI/HPC Systems Engineer to build, deploy, and operate the compute infrastructure supporting high-performance computing and AI development ...
AI/HPC Systems Engineer
San Jose, CA · On-site
$150 - $200/hr
AI/HPC Systems Engineer Position Overview We are seeking an AI/HPC Systems Engineer to build, deploy, and operate the compute infrastructure supporting high-performance computing and AI development ...
AI/HPC Systems Engineer
San Jose, CA · On-site
$78 - $90/hr
AI/HPC Systems Engineer Position Overview We are seeking an AI/HPC Systems Engineer to build, deploy, and operate the compute infrastructure supporting high-performance computing and AI development ...
AI/HPC Systems Engineer
San Jose, CA · On-site
$78 - $90/hr
AI/HPC Systems Engineer Position Overview We are seeking an AI/HPC Systems Engineer to build, deploy, and operate the compute infrastructure supporting high-performance computing and AI development ...
HPC Systems Engineer
Milpitas, CA · On-site
As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex ...
HPC Systems Engineer
Milpitas, CA · On-site
As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex ...
HPC Systems Engineer
Milpitas, CA · On-site
As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex ...
HPC Systems Engineer
Milpitas, CA · On-site
As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex ...
HPC Systems Engineer
Milpitas, CA · On-site
As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex ...
HPC Systems Engineer
Milpitas, CA · On-site
As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex ...
The team operates in an on-premises HPC environment and partners closely with engineering, analysis, IT, security, and specialized vendor partners. About the Job Anduril is looking for an HPC Systems ...
The team operates in an on-premises HPC environment and partners closely with engineering, analysis, IT, security, and specialized vendor partners. About the Job Anduril is looking for an HPC Systems ...
Sr. High Performance Computing (HPC) Systems Engineer
Hawthorne, CA · On-site
$165K - $230K/yr
SR. HIGH PERFORMANCE COMPUTING (HPC) SYSTEMS ENGINEER SpaceX is looking for an HPC Systems Engineer with strong knowledge and experience in a world class engineering organization. This employee will ...
Sr. High Performance Computing (HPC) Systems Engineer
Hawthorne, CA · On-site
$165K - $230K/yr
SR. HIGH PERFORMANCE COMPUTING (HPC) SYSTEMS ENGINEER SpaceX is looking for an HPC Systems Engineer with strong knowledge and experience in a world class engineering organization. This employee will ...
The team operates in an on-premises HPC environment and partners closely with engineering, analysis, IT, security, and specialized vendor partners. About the Job Anduril is looking for an HPC Systems ...
The team operates in an on-premises HPC environment and partners closely with engineering, analysis, IT, security, and specialized vendor partners. About the Job Anduril is looking for an HPC Systems ...
HPC Systems Engineer, Modeling & Simulation
Costa Mesa, CA · On-site
$150 - $200/hr
About the Job Anduril is looking for an HPC Systems Administrator to support the computing ... Partner with analysts, engineers, security teams, and vendor partners to improve reliability ...
HPC Systems Engineer, Modeling & Simulation
Costa Mesa, CA · On-site
$150 - $200/hr
About the Job Anduril is looking for an HPC Systems Administrator to support the computing ... Partner with analysts, engineers, security teams, and vendor partners to improve reliability ...
Sr. High Performance Computing (HPC) Systems Engineer
Hawthorne, CA · On-site
$165K - $230K/yr
Administer and manage HPC clusters, storage systems, and high-speed networks * Provide application support to SpaceX employees across engineering disciplines * Install and integrate Linux-based ...
Sr. High Performance Computing (HPC) Systems Engineer
Hawthorne, CA · On-site
$165K - $230K/yr
Administer and manage HPC clusters, storage systems, and high-speed networks * Provide application support to SpaceX employees across engineering disciplines * Install and integrate Linux-based ...
Sr. High Performance Computing (HPC) Systems Engineer
Hawthorne, CA · On-site
$165K - $230K/yr
Administer and manage HPC clusters, storage systems, and high-speed networks * Provide application support to SpaceX employees across engineering disciplines * Install and integrate Linux-based ...
Sr. High Performance Computing (HPC) Systems Engineer
Hawthorne, CA · On-site
$165K - $230K/yr
Administer and manage HPC clusters, storage systems, and high-speed networks * Provide application support to SpaceX employees across engineering disciplines * Install and integrate Linux-based ...
Sr. High Performance Computing (HPC) Systems Engineer
Hawthorne, CA · On-site
$150 - $200/hr
Sr. High Performance Computing (HPC) Systems Engineer Hawthorne, CA SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one ...
Sr. High Performance Computing (HPC) Systems Engineer
Hawthorne, CA · On-site
$150 - $200/hr
Sr. High Performance Computing (HPC) Systems Engineer Hawthorne, CA SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one ...
AI Systems Engineer - HPC
San Jose, CA · On-site
$151K/yr
The AI Systems Engineer is responsible for the design, development, and administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON: You ...
AI Systems Engineer - HPC
San Jose, CA · On-site
$151K/yr
The AI Systems Engineer is responsible for the design, development, and administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON: You ...
The team operates in an on-premises HPC environment and partners closely with engineering, analysis, IT, security, and specialized vendor partners. About the Job Anduril is looking for an HPC Systems ...
The team operates in an on-premises HPC environment and partners closely with engineering, analysis, IT, security, and specialized vendor partners. About the Job Anduril is looking for an HPC Systems ...
AI Systems Engineer - HPC
San Jose, CA · On-site
The AI Systems Engineer is responsible for the design, development, and administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON: You ...
AI Systems Engineer - HPC
San Jose, CA · On-site
The AI Systems Engineer is responsible for the design, development, and administration of High-Performance Computing (HPC) infrastructure, GPU clusters, and AI workload schedulers. THE PERSON: You ...
AI Systems Engineer - HPC
San Jose, CA · On-site
$150 - $200/hr
... Systems Engineer to join our AMD ITcomputeplatforms engineering team. The AI Systems Engineeris responsible forthe design, development, and administration of High-Performance Computing (HPC ...
AI Systems Engineer - HPC
San Jose, CA · On-site
$150 - $200/hr
... Systems Engineer to join our AMD ITcomputeplatforms engineering team. The AI Systems Engineeris responsible forthe design, development, and administration of High-Performance Computing (HPC ...
AI/HPC System Engineer
San Jose, CA · On-site
$80 - $90/hr
AI/HPC System Engineer Position Description : Protingent Staffing has an exciting contract AI/HPC System Engineer with our client located in San Jose, CA. * We are hiring an AI/HPC System Engineer to ...
AI/HPC System Engineer
San Jose, CA · On-site
$80 - $90/hr
AI/HPC System Engineer Position Description : Protingent Staffing has an exciting contract AI/HPC System Engineer with our client located in San Jose, CA. * We are hiring an AI/HPC System Engineer to ...
AI/HPC System Engineer
San Jose, CA · On-site
AI/HPC System Engineer Location: San Jose, CA (Onsite) Description We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development ...
AI/HPC System Engineer
San Jose, CA · On-site
AI/HPC System Engineer Location: San Jose, CA (Onsite) Description We are hiring an AI/HPC System Engineer to build and operate the compute infrastructure supporting our HPC and AI development ...
Hpc Systems Engineer information
What is an HPC systems engineer?
What are the key skills and qualifications needed to thrive as an HPC systems engineer?
What are some common challenges an HPC systems engineer faces when supporting large-scale computing clusters?
What is the difference between Hpc Systems Engineer vs Hpc Network Engineer?
| Aspect | Hpc Systems Engineer | Hpc Network Engineer |
|---|---|---|
| Credentials | Typically requires a degree in computer science, engineering, or related field; certifications like Cisco CCNA or Linux certifications are common | Similar credentials; often holds networking certifications such as Cisco CCNP or CompTIA Network+ |
| Work Environment | Works on high-performance computing systems, hardware, and software integration in research or enterprise data centers | Focuses on designing, implementing, and maintaining HPC network infrastructure within data centers or research facilities |
| Industry Usage | Used in scientific research, academia, and enterprise sectors with HPC needs | Common in data centers, research institutions, and organizations requiring advanced network performance |
Hpc Systems Engineers and Hpc Network Engineers share overlapping skills in hardware, software, and certifications. However, Hpc Systems Engineers focus on overall system setup and management, while Hpc Network Engineers specialize in network infrastructure. Both roles are vital in supporting high-performance computing environments.
What cities in California are hiring for Hpc Systems Engineer jobs?
Cities in California with the most Hpc Systems Engineer job openings:

HPC Systems Engineer
San Francisco, CA
Full-time
Re-posted 5 days ago
University Of California San Francisco rating
7.8
Based on 13 frontline employees who took The Breakroom Quiz
236th of 633 rated colleges and universities
Job description
The CoreHPC team at UCSF is seeking an HPC Systems Engineer to play a key role in the development, maintenance, and day-to-day operations of the Institute's HPC clusters.
The HPC Systems Engineer will:
- Apply advanced systems infrastructure concepts and skills to the operations and improvement of large-scale and highly complex research Cyber Infrastructure (CI) with unique computing, networking, and storage systems designed to address cutting-edge research problems
- Apply their engineering and design skills to develop new CI solutions, to develop and enhance monitoring to maintain the integrity of CI systems.
- Select methods, techniques and evaluation criteria to develop new CI solutions to address complex research problems.
- Be an active member of the support and maintenance efforts for the CoreHPC cluster, resolving user issues, fixing technical problems, resolving outages, patching, and maintaining systems' uptime and availability.
- Provides consultation, support, and guidance to researchers on how to address computational problems using standard tools, packages, and approaches.
- Develop enhancements of monitoring to maintain the integrity of CI systems.
- Participate in multiple technical projects simultaneously.
- Applies working knowledge of security control frameworks to maintain the integrity of the CI systems and the research being performed on them.
- Gives presentations to the associated team and other technical units.
Evaluates new technologies, including performing moderate to complex cost/benefit analyses.
This position may lead to cross-functional technical working groups and projects in support of onboarding research customers, or making systems improvements.
Department Overview
Academic Research Systems (ARS) serves the needs of the UCSF research community by providing an integrated repository of HIPAA-compliant clinical and life sciences data and a centralized, secure, professionally managed infrastructure for the storage and management of research data. ARS empowers medical scientific investigations by offering secure computing environments, data capture, management and analysis tools, and support services which meet researchers' needs.
The Core HPC team of the Academic Research Service (ARS) focuses on large-scale, high-performance computational and storage services for UCSF researchers so they can address complex computational, AI, and data science problems.
REQUIRED QUALIFICATIONS
- Bachelor's degree in a related area such as computer science or engineering, and 6+ years of experience with large-scale or HPC systems * or* 10+ years of related experience with large-scale or HPC systems
- Expert knowledge of HPC systems infrastructure design
- Strong knowledge of high-performance parallel filesystems and storage such as GPFS, Lustre, Vast, DDN, etc.
- Advanced knowledge of computer security best practices and policies including demonstrated experience securing research cyberinfrastructure systems to meet NIST 800-171 / 800-223, HIPPA or IS-3 requirements
- Demonstrated testing and test planning skills. Demonstrated ability to create automated testing.
- Knowledge of HPC job scheduler system design and operation such as SLURM or PBS,
- Demonstrated skill (5 years +) deploying, managing, and troubleshooting Warewulf (or similar) infiniband based clusters
- Ability to elicit and communicate technical and non-technical information in a clear and concise manner.
- Self-motivated and works independently and as part of a team. Demonstrates problem-solving skills. Able to learn effectively and meet deadlines.
- Understanding of system performance monitoring and actions that can be taken to improve or correct performance.
- Demonstrated advanced knowledge, skills and abilities associated with system problem identification and resolution. Experience with design, configuration, operation, repair, and tuning of technology systems.
- Advanced experience writing and editing the most complex scripts used to perform system maintenance and administration.
Ability to write technical documentation in a clear and concise manner. Ability to develop runbooks defining complex technical processes in a clear and concise manner
PREFERRED QUALIFICATIONS
- Knowledge of the design, development, and application of technology and systems to meet business needs.
- General knowledge of other areas of IT. Thorough understanding of and experience with systems-related issues and actions that can be taken to improve or correct performance.
- Demonstrated skills associated with adapting equipment and technology to serve user needs. Demonstrated comprehensive understanding of how system management actions affect other systems, system users and dependent/related functions.
%
of time
Essential Function (Yes/No)
Key Responsibilities
(To be completed by Supervisor)
15
Applies advanced systems / infrastructure concepts to define, design, implement, and operate highly complex, research cyberinfrastructure systems, services and technology solutions. Proposes and implements highly complex system or device enhancements such as software, hardware and network configuration, updates and installations for projects or services of broad scope. Sets standards for monitoring and maintaining the health and integrity of CI systems including upgrading and patching.15
Independently manages systems and services for a large facility, campuswide, medical center or Office of the President and / or institution-wide scope and makes recommendations for purchases or upgrades. Performs complex and advanced analysis to acquire, install, modify and support operating systems, databases, utilities and web-related tools. Selects methods and techniques to obtain solutions. Interacts with senior management. May perform complex network integration tasks and interoperability assessments for interconnected servers or components of clusters for communication. Support and collaborate with researchers and other key IT (e.g. network and security) and Data Center partners in a timely manner15
Specifies, writes and executes highly complex software and scripts to support systems management, log analysis, monitoring, deployment, configuration management, and other system administration duties for multiple, highly integrated systems.30
Provides consultation, training, support, and guidance to researchers enabling them to utilize HPC resources effectively.10
Maintains complex security systems. Interprets and adopts campus, medical center or Office of the President, system and regulation-based security policies to control access to networked resources. Provides recommendations and requirements on network access controls.5
Collaborates and may provide leadership with other Systems Engineers within the CI ecosystem/higher-education community. Regularly contribute best practices documentation, present at conferences, or publish in peer reviewed journals.10
Define and track performance metrics to ensure efficient current and future use of cyber infrastructure resources.100%
(To update total %, enter the amount of time in whole numbers (without the % symbol - e.g., 15, 20) then highlight the total sum (e.g., 1%) at the bottom of the column and press F9. The total sum should add up to 100%.)What University Of California San Francisco employees say
Pay
Benefits
Hours and flexibility
Workplace
Get the full story on Breakroom
About University of California San Francisco
Sourced by ZipRecruiter
Industry
Colleges, universities, and professional schools
Company size
10,000+ Employees
Headquarters location
San Francisco, CA, US