As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale ...
As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale ...
Cloud HPC Engineer
Walnut Creek, CA · On-site +1
$130K - $200K/yr
Work will focus on (1) building a cloud-based high-performance computing platform for nanoscale materials and chemistry, (2) planning and organizing the work of the engineering team, (3) establishing ...
Cloud HPC Engineer
Walnut Creek, CA · On-site +1
$130K - $200K/yr
Work will focus on (1) building a cloud-based high-performance computing platform for nanoscale materials and chemistry, (2) planning and organizing the work of the engineering team, (3) establishing ...
HPC Systems Engineer
Chicago, IL · On-site
Radix Trading is seeking a seasoned HPC Systems Engineer with a passion for Linux, HPC systems and ... Qualifications: * 5+ years of HPC system administration/architecture including RHEL/CentOS/Rocky ...
HPC Systems Engineer
Chicago, IL · On-site
Radix Trading is seeking a seasoned HPC Systems Engineer with a passion for Linux, HPC systems and ... Qualifications: * 5+ years of HPC system administration/architecture including RHEL/CentOS/Rocky ...
HPC Systems Engineer
Chicago, IL · On-site
Radix Trading is seeking a seasoned HPC Systems Engineer with a passion for Linux, HPC systems and ... Qualifications: * 5+ years of HPC system administration/architecture including RHEL/CentOS/Rocky ...
HPC Systems Engineer
Chicago, IL · On-site
Radix Trading is seeking a seasoned HPC Systems Engineer with a passion for Linux, HPC systems and ... Qualifications: * 5+ years of HPC system administration/architecture including RHEL/CentOS/Rocky ...
Sr. HPC Cloud Developer Lead
Mountain View, CA · Hybrid
$66.25 - $90.75/hr
We offer services ranging from full life cycle HPC systems engineering to remote managed services ... We are seeking a AWS Cloud SME to support a cloud contract supporting the National Environmental ...
Sr. HPC Cloud Developer Lead
Mountain View, CA · Hybrid
$66.25 - $90.75/hr
We offer services ranging from full life cycle HPC systems engineering to remote managed services ... We are seeking a AWS Cloud SME to support a cloud contract supporting the National Environmental ...
Senior HPC System Administrator
Albuquerque, NM · On-site
$125 - $150/hr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
Senior HPC System Administrator
Albuquerque, NM · On-site
$125 - $150/hr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
HPC System Administrator - Hybrid
Irving, TX · On-site
$125 - $150/hr
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical field, or three (3) years of directly related Linux systems administration or HPC infrastructure ...
HPC System Administrator - Hybrid
Irving, TX · On-site
$125 - $150/hr
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical field, or three (3) years of directly related Linux systems administration or HPC infrastructure ...
AI/HPC System Performance Engineer
Menlo Park, CA · On-site
$154K - $217K/yr
As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale ...
AI/HPC System Performance Engineer
Menlo Park, CA · On-site
$154K - $217K/yr
As an AI/HPC System Performance Engineer on the Network Infrastructure Engineering team, you will drive end-to-end performance characterization, bottleneck analysis, and optimization of large-scale ...
Senior HPC Systems Architect
San Jose, CA · On-site
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
Senior HPC Systems Architect
San Jose, CA · On-site
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
Senior HPC System Administrator
Albuquerque, NM · On-site
$125 - $150/hr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
Senior HPC System Administrator
Albuquerque, NM · On-site
$125 - $150/hr
Work with HPC team and vendors in the delivery and deployment of large storage systems, to include facilities staff and hardware engineers. * Actively seek opportunities to broaden and deepen your ...
... programming models (e.g., MPI, OpenMP) • Strong knowledge of HPC system architectures and ... cloud-based HPC platforms and hybrid solutions • Understanding of emerging HPC technologies (e.g ...
... programming models (e.g., MPI, OpenMP) • Strong knowledge of HPC system architectures and ... cloud-based HPC platforms and hybrid solutions • Understanding of emerging HPC technologies (e.g ...
AI/HPC Systems Engineer
San Jose, CA · On-site
$150 - $200/hr
Monitor system performance, availability, and resource utilization to ensure reliable operations ... Hybrid Cloud Infrastructure Deploy and maintain computing environments across on-premises ...
AI/HPC Systems Engineer
San Jose, CA · On-site
$150 - $200/hr
Monitor system performance, availability, and resource utilization to ensure reliable operations ... Hybrid Cloud Infrastructure Deploy and maintain computing environments across on-premises ...
Deploying and configuring HPC software (operating systems, job schedulers, parallel programming environments) * Setting up network infrastructure for high-speed data transfer * Performance ...
Deploying and configuring HPC software (operating systems, job schedulers, parallel programming environments) * Setting up network infrastructure for high-speed data transfer * Performance ...
Troubleshoot application-related issues and ensure minimal disruption to engineering activities ... Perform system health checks, monitoring, and incident tracking for HPC and CAE environments.
Quick apply
Troubleshoot application-related issues and ensure minimal disruption to engineering activities ... Perform system health checks, monitoring, and incident tracking for HPC and CAE environments.
Senior HPC Systems Architect
San Jose, CA · On-site
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
Senior HPC Systems Architect
San Jose, CA · On-site
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
Senior HPC Systems Architect
San Jose, CA · On-site
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
Senior HPC Systems Architect
San Jose, CA · On-site
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
Deploying and configuring HPC software (operating systems, job schedulers, parallel programming environments) * Setting up network infrastructure for high-speed data transfer * Performance ...
Deploying and configuring HPC software (operating systems, job schedulers, parallel programming environments) * Setting up network infrastructure for high-speed data transfer * Performance ...
Senior HPC Systems Architect
San Francisco, CA · On-site
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
Senior HPC Systems Architect
San Francisco, CA · On-site
$200 - $250/hr
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of ... Our scope spans from rack and pod level system arrangement that maximizes the data center power ...
HPC Software Architect
San Antonio, TX · On-site
... programming models (e.g., MPI, OpenMP) • Strong knowledge of HPC system architectures and ... cloud-based HPC platforms and hybrid solutions • Understanding of emerging HPC technologies (e.g ...
HPC Software Architect
San Antonio, TX · On-site
... programming models (e.g., MPI, OpenMP) • Strong knowledge of HPC system architectures and ... cloud-based HPC platforms and hybrid solutions • Understanding of emerging HPC technologies (e.g ...
Cloud Hpc System Engineer information
See salary details
$23.56 - $29.35
0% of jobs
$29.35 - $35.14
1% of jobs
$35.14 - $40.93
2% of jobs
$40.93 - $46.72
6% of jobs
$46.72 - $52.51
12% of jobs
$53.72 is the 25th percentile. Wages below this are outliers.
$52.51 - $58.30
19% of jobs
The median wage is $61.36 / hr.
$58.30 - $64.10
19% of jobs
$64.10 - $69.89
16% of jobs
$70.02 is the 75th percentile. Wages above this are outliers.
$69.89 - $75.68
12% of jobs
$75.68 - $81.47
7% of jobs
$81.47 - $87.26
6% of jobs
$23
$62
$87
How much do cloud hpc system engineer jobs pay per hour?
What are popular job titles related to Cloud Hpc System Engineer jobs?
For Cloud Hpc System Engineer jobs, the most frequently searched job titles are:
AI/HPC System Performance Engineer
Menlo Park, CA
$154K/yr
Full-time
Posted 27 days ago
Meta rating
7.8
Based on 45 frontline employees who took The Breakroom Quiz
139th of 247 rated software companies
Job description
AI/HPC System Performance Engineer Responsibilities:
- Profile and benchmark AI training and inference workloads across large-scale HPC clusters to identify network, compute, and memory bottlenecks
- Develop and maintain performance analysis frameworks and dashboards to track system-level metrics including GPU utilization, network bandwidth, latency, and collective communication efficiency
- Investigate and resolve performance regressions in distributed AI training environments, including issues related to RDMA fabrics, collective communication libraries, and job scheduling
- Collaborate with network infrastructure, hardware, and AI research teams to define performance requirements and validate new HPC cluster configurations
- Design and execute capacity and scalability experiments to inform network topology decisions for AI supercomputing infrastructure
- Build tooling and automation to continuously monitor HPC system health, detect anomalies, and reduce mean time to mitigation during performance incidents
- Establish service level objectives for AI cluster network performance and drive cross-functional alignment on reliability and efficiency targets
- Lead technical design reviews for network and system architecture changes affecting AI workload performance, communicating trade-offs clearly to engineering and product stakeholders
- Mentor other engineers on HPC performance methodologies, debugging techniques, and instrumentation best practices
- Leverage AI-assisted workflows to accelerate root cause analysis, automate routine performance reporting, and expand coverage across the HPC stack
Minimum Qualifications:
- 5+ years of coding experience in C, C++, Python, or similar programming languages. Flexible to learn new programming languages
- Experience profiling and optimizing distributed AI or HPC workloads, including familiarity with GPU interconnects, RDMA networking, and collective communication frameworks such as NCCL or MPI
- Experience debugging complex, non-reproducible performance issues across multi-layer systems including network fabric, operating system, and application layers
- Experience designing and implementing performance monitoring systems, including instrumentation, telemetry pipelines, and alerting for large-scale infrastructure
- Experience driving cross-functional technical projects from requirements definition through production deployment, including communicating performance findings and trade-offs to diverse stakeholders
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- 6+ years of experience in system performance engineering, network infrastructure engineering, or a related field within large-scale distributed computing or HPC environments
Preferred Qualifications:
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Understanding of RDMA congestion control mechanisms on IB and RoCE Networks
- Understanding of AI training workloads and demands they exert on networks
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Understanding of the latest artificial intelligence (AI) technologies
- Experience in developing systems software in languages like C++
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Experience with machine learning frameworks such as PyTorch and TensorFlow
About Meta:
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.
$154,000/year to $217,000/year + bonus + equity + benefits
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.
About Meta
Sourced by ZipRecruiter
Industry
Internet and it, media and telecom and software development
Company size
10,000+ Employees
Headquarters location
Menlo Park, CA, US