1

Hpc Performance Engineer Jobs in Washington (NOW HIRING)

Responsibilities : • Design and develop software for high-performance computing systems • ... programming languages (e.g., MPI, OpenMP) • Strong understanding of HPC architectures and ...

Citizenship) KEY SUMMARY We are seeking an innovative and driven HPC (High-Performance Computing) Software Engineer to develop and optimize software solutions for cutting-edge computational ...

Citizenship) KEY SUMMARY We are seeking an innovative and driven HPC (High-Performance Computing) Software Engineer to develop and optimize software solutions for cutting-edge computational ...

Citizenship) KEY SUMMARY We are seeking an innovative and driven HPC (High-Performance Computing) Software Engineer to develop and optimize software solutions for cutting-edge computational ...

Citizenship) KEY SUMMARY We are seeking an innovative and driven HPC (High-Performance Computing) Software Engineer to develop and optimize software solutions for cutting-edge computational ...

Citizenship) KEY SUMMARY We are seeking an innovative and driven HPC (High-Performance Computing) Software Engineer to develop and optimize software solutions for cutting-edge computational ...

HPC Engineer

Rockville, MD · Hybrid

$123K - $166K/yr

High Performance Computing (HPC), High-Performance Computing (HPC) Systems, Scientific Software ... No We are seeking an HPC Engineer to join a Scientific Computing team supporting high-performance ...

Showing results 21-40

Hpc Performance Engineer information

What is an HPC performance engineer?

HPC Performance Engineers are specialists who focus on optimizing the performance of high-performance computing (HPC) systems and applications. They analyze system bottlenecks, tune software and hardware configurations, and work with researchers and developers to ensure applications run efficiently on supercomputers or large computing clusters. Their work is essential for maximizing computational resources and improving the speed and scalability of scientific, engineering, or data-intensive workloads.

What are the key skills and qualifications needed to thrive as an HPC performance engineer?

To thrive as an HPC Performance Engineer, you need a strong background in computer science or engineering, with expertise in parallel programming, high-performance computing architectures, and performance analysis. Familiarity with tools like MPI, OpenMP, profiling software (e.g., Intel VTune, GNU gprof), and experience with job schedulers and Linux systems are essential. Analytical thinking, problem-solving, and effective communication are crucial soft skills for identifying bottlenecks and collaborating with multidisciplinary teams. These skills are vital for optimizing computational workflows, maximizing resource utilization, and driving efficiency in complex HPC environments.

What are the typical challenges HPC performance engineers face when optimizing large-scale computational workloads?

HPC Performance Engineers often encounter challenges such as identifying bottlenecks in parallel code, managing resource contention, and optimizing data movement across distributed systems. They must balance maximizing throughput with minimizing latency, all while ensuring applications scale efficiently as cluster sizes grow. Collaboration with software developers, system administrators, and research teams is common to align application requirements with hardware capabilities and to implement effective performance improvements.

What is the difference between Hpc Performance Engineer vs Hpc System Administrator?

AspectHpc Performance EngineerHpc System Administrator
Primary FocusOptimizing HPC system performance and efficiencyManaging and maintaining HPC infrastructure
Skills & CertificationsPerformance tuning, parallel computing, Linux, scriptingSystem setup, network management, user support
Work EnvironmentResearch labs, data centers, high-performance computing facilitiesData centers, IT departments, research institutions
Common TasksPerformance analysis, bottleneck resolution, code optimizationSystem installation, user account management, hardware troubleshooting

The Hpc Performance Engineer focuses on enhancing system performance and efficiency, often working on optimization and tuning. In contrast, the Hpc System Administrator manages the day-to-day operation and maintenance of HPC systems. Both roles are essential in high-performance computing environments but serve different core functions.

External Job Posting Title Senior HPC DevOps Engineer

College Park, MD • On-site

Peraton
IT Services • 10K+ employees

$128K - $165K/yr

Other

Posted 12 days ago


Key responsibilities

  • Own and manage automation workflows for HPC/AI compute clusters, including job templates, inventories, credentials, RBAC configurations, and execution environments.

  • Implement desired-state enforcement, drift detection, and alerting to maintain cluster service configurations and reconcile runtime versus configured states.

  • Build and maintain automated workflows for onboarding compute nodes, including OS installation, security baselines, hardware validation, and reporting pass/fail results.


Peraton rating

8.3

Company rating: 8.3 out of 10

Based on 56 frontline employees who took The Breakroom Quiz


Job description

Responsibilities

Peraton Labs is seeking a poly cleared Senior HPC DevOps Engineer to own the operations and automation lifecycle for an existing HPC/AI compute cluster (Linux). You will work closely with Peraton team members, as well as directly with our Maryland-based customer, in a fast-paced environment at a customer site. In this role you will codify repeatable operations in Ansible and drive execution through an enterprise automation controller to enforce desired state, detect drift, accelerate node onboarding, and streamline incident response via runbook automation integrated with monitoring and ITSM.

Full-time on-site work at a customer site near College Park, MD.

Key responsibilities may include

  • Automation ownership: Own and manage automation workflows, including job templates, inventories, credentials, RBAC configurations, execution environments, and promotion across environments.
  • Desired-state and drift detection: Enforce desired state across cluster services via code-driven configuration; implement drift detection and alert on deviations; reconcile runtime state vs configured state.
  • Compute node onboarding (Bare-metal/VM): Build and maintain an automated node bootstrap workflow that installs/configures the OS, applies security and performance baselines, enrolls nodes into the scheduler and shared storage ecosystem, validates hardware and service readiness (CPU, network, accelerator, storage mounts), and reports pass/fail results.
  • Patch & vulnerability response: Implement rolling maintenance and patch automation to meet defined vulnerability response SLAs. Maintain version-controlled container build definitions and integrate image scanning into the build/release lifecycle.
  • Logging & observability: Ensure automation and operational workflows emit auditable logs to centralized analytics and integrate with metrics/alerting to enable reliable incident response, proactive detection, and safe auto-remediation.
  • Incident/problem management: Automate responses to common incidents (hung nodes, storage performance alarms, image vulnerabilities, hardware failures) leveraging out-of-band hardware management interfaces and standardized runbooks.
  • Docs-as-code: Keep runbooks and operational documentation versioned alongside automation and publish operator guidance to the orgs documentation platform.

This position may be eligible for an increased sign-on bonus. Eligibility, bonus amount, and applicable terms and conditions will be discussed during the recruiting process

#MDFSP

#PLABS26

Qualifications

Required qualifications

  • 12+ years of experience and a BS in computer science, IT, or related technical field, MS and 10 years of experience, or a Ph.D. with 8 years of experience. Four years of additional experience is required in lieu of a Bachelors’ degree for a total of 16 years of experience.
  • 7+ years in Linux systems / SRE / DevOps, including production cluster operations in an HPC or large-scale compute environment.
  • 3+ years of experience building and operating Ansible automation at scale (roles/collections, idempotency, inventories, secrets).
  • Strong Linux hardening & compliance fundamentals (SELinux/AppArmor, SSH key automation, baseline config management).
  • Demonstrated experience operating or automating clustered compute environments (HPC, large Linux farms, or similar).
  • Hands-on experience with container tooling in Linux environments, including image lifecycle/versioning.
  • Familiarity with incident response and runbook-driven operations; ability to automate common remediations.
  • Strong Git workflow and documentation practices.
  • Must hold at least one active/current technical certification from the following-
    • Systems engineering (e.g., INCOSE)
    • Information security (e.g., CISSP)
    • Networking (e.g., CCNA)
    • System Administration (e.g., RHCE, MCSE)
    • Virtualization (e.g., VCP)
    • IT systems management (e.g., ITIL)
    • Project management (e.g., PMP, Agile)
  • Active TS/SCI security clearance with a current polygraph is required

Preferred qualifications

  • Bare-metal provisioning experience (PXE/iPXE, Kickstart/Preseed, Foreman/MAAS) and hardware OOB management.
  • CI/testing for automation and promotion pipelines for playbooks
  • Experience with tuned performance profiles, HPC performance troubleshooting, and GPU node health validation.

#MDPM

#MDFSP

Peraton Overview

Peraton is a next-generation national security company that drives missions of consequence spanning the globe and extending to the farthest reaches of the galaxy. As the world's leading mission capability integrator and transformative enterprise IT provider, we deliver trusted, highly differentiated solutions and technologies to protect our nation and allies. Peraton operates at the critical nexus between traditional and nontraditional threats across all domains: land, sea, space, air, and cyberspace. The company serves as a valued partner to essential government agencies and supports every branch of the U.S. armed forces. Each day, our employees do the can't be done by solving the most daunting challenges facing our customers. Visit peraton.com to learn how we're keeping people around the world safe and secure.

Target Salary Range

$146,000 - $234,000. This represents the typical salary range for this position. Salary is determined by various factors, including but not limited to, the scope and responsibilities of the position, the individual’s experience, education, knowledge, skills, and competencies, as well as geographic location and business and contract considerations. Depending on the position, employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.

EEO

EEO: Equal opportunity employer, including disability and protected veterans, or other characteristics protected by law.

#J-18808-Ljbffr

What Peraton employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Peraton logo

About Peraton

Sourced by ZipRecruiter

At Peraton, we re at the forefront of delivering the next big thing every day. We re the partner of choice to help solve some of the world s most daunting challenges, delivering bold, new solutions to keep people around the world safer and more secure.

Industry

It services

Company size

10,000+ Employees

Headquarters location

Herndon, VA, US

Year founded

2017