Senior HPC Specialist
Denver, CO · On-site
... e.g., Slurm, PBS, Grid Engine). • Monitor system health, troubleshoot issues, and resolve performance bottlenecks. • Ensure optimal configuration and high availability of HPC resources. • ...
Denver, CO · On-site
... e.g., Slurm, PBS, Grid Engine). • Monitor system health, troubleshoot issues, and resolve performance bottlenecks. • Ensure optimal configuration and high availability of HPC resources. • ...
Denver, CO · On-site
... e.g., Slurm, PBS, Grid Engine). • Monitor system health, troubleshoot issues, and resolve performance bottlenecks. • Ensure optimal configuration and high availability of HPC resources. • ...
Thornton, CO · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Thornton, CO · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Greeley, CO · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Greeley, CO · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Aurora, CO · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Aurora, CO · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Denver, CO · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Denver, CO · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
PBSPro, SLURM) • Experience with Panasas/Vdura storage support • Proficiency in technical writing and documentation of solutions • Demonstrated ability to work seamlessly across organization ...
PBSPro, SLURM) • Experience with Panasas/Vdura storage support • Proficiency in technical writing and documentation of solutions • Demonstrated ability to work seamlessly across organization ...
Manage job scheduling and workload optimization using tools like SLURM. * Administer parallel file systems (such as ceph and IBM Spectrum Scale/GPFS) and storage solutions. Development * Design and ...
Manage job scheduling and workload optimization using tools like SLURM. * Administer parallel file systems (such as ceph and IBM Spectrum Scale/GPFS) and storage solutions. Development * Design and ...
Manage job scheduling and workload optimization using tools like SLURM. * Administer parallel file systems (such as ceph and IBM Spectrum Scale/GPFS) and storage solutions. Development * Design and ...
Manage job scheduling and workload optimization using tools like SLURM. * Administer parallel file systems (such as ceph and IBM Spectrum Scale/GPFS) and storage solutions. Development * Design and ...
Broomfield, CO · On-site
$132K - $165K/yr
Develop robust interfaces between HPC job schedulers (e.g., Slurm) and quantum runtimes, optimizing data movement and job scheduling. * Work with QPU hardware control systems and SDKs to integrate ...
Broomfield, CO · On-site
$132K - $165K/yr
Develop robust interfaces between HPC job schedulers (e.g., Slurm) and quantum runtimes, optimizing data movement and job scheduling. * Work with QPU hardware control systems and SDKs to integrate ...
Centennial, CO · On-site
$154 - $195/hr
Multi-GPU and distributed training experience (PyTorch DDP or FSDP), including shared cluster resources (Slurm, Kubernetes, Run:ai) * Experience with experiment tracking and data/model versioning for ...
Centennial, CO · On-site
$154 - $195/hr
Multi-GPU and distributed training experience (PyTorch DDP or FSDP), including shared cluster resources (Slurm, Kubernetes, Run:ai) * Experience with experiment tracking and data/model versioning for ...
High-performance computing (HPC) systems, supercomputing environments, or HPC job schedulers such as Slurm or PBS. * GPU or accelerator programming, including CUDA or similar programming models.
High-performance computing (HPC) systems, supercomputing environments, or HPC job schedulers such as Slurm or PBS. * GPU or accelerator programming, including CUDA or similar programming models.
Boulder, CO · On-site
$165K - $225K/yr
High-performance computing (HPC) systems, supercomputing environments, or HPC job schedulers such as Slurm or PBS. * GPU or accelerator programming, including CUDA or similar programming models.
Quick apply
Boulder, CO · On-site
$165K - $225K/yr
High-performance computing (HPC) systems, supercomputing environments, or HPC job schedulers such as Slurm or PBS. * GPU or accelerator programming, including CUDA or similar programming models.
For Slurm jobs in Colorado, the most frequently searched job titles are:
The top searched job categories for Slurm jobs in Colorado are:

Sourced by ZipRecruiter
Custom software development services
51 - 200 Employees
Los Altos, CA, US