1

Hpc Cluster Administrator Slurm Jobs (NOW HIRING)

Hands-on experience administering an HPC cluster (workload scheduler such as PBS Professional or ... Administer, patch, and sustain RHEL servers and workstations across three or more isolated ...

Solid understanding of multiple operating systems and cluster technologies. * Experience with ... Understanding of HPC platforms to support users with SLURM job submissions and troubleshooting.

Administer and optimize enterprise storage platforms such as Qumulo and NetApp in support of HPC ... Understanding of HPC job schedulers (SLURM) and user support workflows. * Experience with container ...

Hands-on experience administering an HPC cluster (workload scheduler such as PBS Professional or ... Administer, patch, and sustain RHEL servers and workstations across three or more isolated ...

HPC/ML Infrastructure Engineer

San Francisco, CA · On-site

$126K - $166K/yr

... cluster. • Serve as the bridge between researchers and GPU machines. • Ensure SLURM jobs are ... Required : • Experience in HPC infrastructure engineering • Familiarity with modern HPC ...

Showing results 41-60

Hpc Cluster Administrator Slurm information

See salary details

$14

$36

$63

How much do hpc cluster administrator slurm jobs pay per hour?

As of Aug 24, 2026, the average hourly pay for hpc cluster administrator slurm in the United States is $36.33, according to ZipRecruiter salary data. Most workers in this role earn between $24.28 and $50.96 per hour, depending on experience, location, and employer.

What does an HPC cluster administrator do with Slurm?

An HPC (High Performance Computing) Cluster Administrator who specializes in Slurm is responsible for managing, configuring, and maintaining large-scale computer clusters used for scientific and engineering computations. They install and optimize the Slurm workload manager, which schedules and allocates computing resources to users and jobs efficiently. Their duties include monitoring cluster health, troubleshooting issues, managing user access, and updating software and hardware components. Additionally, they help users with job submissions, ensure security protocols are followed, and work to maximize system uptime and performance.

What are common challenges faced by an HPC cluster administrator working with Slurm, and how can they be addressed?

One common challenge HPC Cluster Administrators encounter is efficiently managing resource allocation and job scheduling in environments with diverse workloads. Balancing user demands, optimizing Slurm configurations, and troubleshooting job failures require both technical expertise and strong communication skills. Proactively monitoring system health, keeping Slurm up-to-date, and collaborating closely with researchers and IT teams help address these challenges. Regular training and staying engaged with the Slurm user community also enable administrators to implement best practices and quickly resolve issues.

What are the key skills and qualifications needed to thrive as an HPC cluster administrator with Slurm?

To thrive as an HPC Cluster Administrator specializing in Slurm, you need expertise in Linux system administration, networking, and parallel computing, often supported by a degree in computer science or a related field. Familiarity with Slurm workload manager, scripting languages (such as Bash or Python), and configuration management tools like Ansible is essential, and certifications like RHCE can be advantageous. Strong problem-solving skills, attention to detail, and effective communication are important soft skills for collaborating with researchers and troubleshooting complex issues. These abilities ensure the efficient operation, reliability, and scalability of high-performance computing environments vital for research and development.

What is the difference between Hpc Cluster Administrator Slurm vs Hpc Cluster Engineer?

AspectHpc Cluster Administrator SlurmHpc Cluster Engineer
CredentialsTypically requires Linux certifications, HPC-specific training, and Slurm knowledgeRequires similar Linux certifications, scripting skills, and HPC system understanding
Work EnvironmentFocuses on managing and maintaining HPC clusters using Slurm workload managerDesigns, develops, and optimizes HPC systems and workflows
Industry UsageCommonly employed in research institutions, universities, and labs using SlurmFound in research, scientific computing, and high-performance computing sectors

Hpc Cluster Administrator Slurm primarily manages and maintains HPC clusters with a focus on Slurm workload management, ensuring system stability and job scheduling. In contrast, Hpc Cluster Engineer designs and develops HPC systems, often working on performance optimization and infrastructure development. Both roles require Linux expertise and HPC knowledge but differ in their core responsibilities and focus areas.

More about Hpc Cluster Administrator Slurm jobs

What cities are hiring for Hpc Cluster Administrator Slurm jobs?

Cities with the most Hpc Cluster Administrator Slurm job openings:

What states have the most Hpc Cluster Administrator Slurm jobs?

States with the most job openings for Hpc Cluster Administrator Slurm jobs include:

What job categories do people searching Hpc Cluster Administrator Slurm jobs look for?

The top searched job categories for Hpc Cluster Administrator Slurm jobs are:

Infographic showing various Hpc Cluster Administrator Slurm job openings in the United States as of August 2026, with employment types broken down into 2% As Needed, 83% Full Time, 10% Part Time, 1% Temporary, and 4% Contract. Highlights an 91% Physical, 3% Hybrid, and 6% Remote job distribution, with an average salary of $75,575 per year, or $36.3 per hour.

Sr. Linux Systems Administrator

Astrion

Columbia, MD • On-site

$145K - $165K/yr

Full-time

Posted 20 days ago


Job description

Overview
Senior Linux Systems Administrator
LOCATION: Columbia, MD (Onsite)
JOB STATUS: Full-time
CLEARANCE: Active Top Secret / TS/SCI (Required)
TRAVEL: Less than 10%
SALARY RANGE: $145k - $165k
Astrion has an exciting opportunity for a highly experienced Senior Linux Systems Administrator to build, secure, and sustain the Linux server and workstation estate, the high-performance computing (HPC) cluster, and the enterprise storage that run across three or more isolated classified networks supporting Department of Defense/Department of War (DoD/DoW) environments in Columbia, MD. The ideal candidate brings deep Red Hat Enterprise Linux (RHEL) expertise, extensive experience operating inside closed and air-gapped classified enclaves, and a proven record of building and maintaining systems in compliance with Risk Management Framework (RMF), DoD STIG, and Comply-to-Connect (C2C) requirements. This position requires Monday through Friday on-site work at our facility in Columbia, MD.
REQUIRED QUALIFICATIONS / SKILLS
  • Active TS/SCI security clearance (required)
  • Bachelor's degree in Computer Science, Information Technology, or related field (or equivalent experience)
  • 8+ years of Linux systems administration experience within DoD/DoW or classified environments
  • Deep expertise in:
    • RHEL administration (installation, patching, performance tuning, kernel and systemd, storage, and networking)
    • Scripting and automation (Bash, Python, Ansible)
    • Linux identity and authentication (LDAP/SSSD, Kerberos, FreeIPA or AD integration)
  • Demonstrated experience implementing and maintaining DoD STIG compliance on Linux
  • Hands-on experience administering an HPC cluster (workload scheduler such as PBS Professional or similar)
  • Experience administering enterprise storage (NetApp ONTAP, NFS)
  • Strong understanding of:
    • RMF (Risk Management Framework)
    • DISA security requirements and accreditation processes
    • DCSA accreditation standards
  • Active DoD 8140 (formerly DoD 8570) compliant baseline certification (e.g., Security+ CE, CISSP)
  • Active DoD 8140 compliant computing environment certification (e.g., minimum Red Hat Certified System Administrator (RHCSA))
  • Experience with:
    • Offline and air-gapped patching (Red Hat Satellite or internal repositories)
    • ACAS or Nessus and SCAP scanning
    • SIEM integration and log analysis
    • Continuous monitoring under RMF

KEY COMPETENCIES
  • Advanced Linux systems administration and troubleshooting
  • Security hardening and compliance enforcement
  • Automation and scripting
  • Strong analytical and problem-solving abilities
  • Excellent communication and documentation skills
  • Ability to operate in high-security, mission-critical environments

PREFERRED QUALIFICATIONS / SKILLS
  • Red Hat Certified Engineer (RHCE) or Red Hat Certified Architect (RHCA)
  • HPC experience with PBS Professional
  • NetApp Certified Data Administrator (NCDA)
  • Automation at scale (Ansible Automation Platform, Git-based configuration management)
  • Zero Trust architectures and containerization (Podman, Kubernetes) in classified environments
  • Prior experience supporting Cross Domain Solutions (CDS) programs
  • Experience with cloud-based DoD environments (e.g., Azure Government, Azure Secret)
  • Project management experience and experience briefing executive leadership

RESPONSIBILITIES
  • Administer, patch, and sustain RHEL servers and workstations across three or more isolated classified networks, each operated as its own authorization boundary
  • Build, harden, and baseline Linux systems to DoD Security Technical Implementation Guides (STIGs); remediate findings and control configuration drift
  • Administer and tune the HPC cluster, including compute nodes, the workload scheduler (e.g., PBS Professional), the shared or parallel filesystem, and the low-latency interconnect
  • Administer NetApp storage (ONTAP), including volumes, NFS and CIFS exports, snapshots, capacity planning, and backup and recovery
  • Automate provisioning, configuration, and remediation with Bash, Python, and Ansible, including on disconnected systems
  • Perform patching and updates on air-gapped networks through approved offline processes (e.g., Red Hat Satellite or internal repositories and trusted media transfer)
  • Administer Linux identity and access (LDAP/SSSD, Kerberos, FreeIPA or Active Directory integration, sudo, PAM), enforcing least privilege
  • Implement and maintain C2C and 802.1x posture for Linux endpoints and integrate with the security tool stack
  • Perform vulnerability remediation and continuous monitoring in accordance with RMF controls; run and interpret ACAS and SCAP scans and review audit logs
  • Support the Authority to Operate (ATO) process, including STIG checklists, POA&Ms, and risk assessments
  • Develop and maintain system documentation, diagrams, SOPs, and security artifacts
  • Troubleshoot complex issues across multi-network, multi-vendor environments
  • Collaborate with cybersecurity, network, systems engineering, and mission stakeholders to ensure secure, reliable operations
  • Support audits, inspections, and compliance validation activities