Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Columbus, OH · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Columbus, OH · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Toledo, OH · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Toledo, OH · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Cleveland, OH · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Cleveland, OH · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Setting up, configuring, and continuously evolving NEURA's HyperPod clusters, including HyperPod/Slurm and HyperPod/EKS orchestration models. * Designing and implementing strategies for cluster ...
Setting up, configuring, and continuously evolving NEURA's HyperPod clusters, including HyperPod/Slurm and HyperPod/EKS orchestration models. * Designing and implementing strategies for cluster ...
Lead integration ofcompute, storage, networking, the AI software stack (CUDA,ROCm, Triton, NIM, NVIDIA AI Enterprise,Run:ai,Slurm, Kubernetes / Kubeflow) and managed-service operating models across ...
Lead integration ofcompute, storage, networking, the AI software stack (CUDA,ROCm, Triton, NIM, NVIDIA AI Enterprise,Run:ai,Slurm, Kubernetes / Kubeflow) and managed-service operating models across ...
Junior Computer Operator
Dayton, OH · On-site
$15.75 - $19.25/hr
Continuously monitor the health, availability, and utilization of HPC clusters using system dashboards and workload management software (such as Slurm). * Facility Oversight: Monitor critical data ...
Junior Computer Operator
Dayton, OH · On-site
$15.75 - $19.25/hr
Continuously monitor the health, availability, and utilization of HPC clusters using system dashboards and workload management software (such as Slurm). * Facility Oversight: Monitor critical data ...
Junior Computer Operator
Dayton, OH · On-site
$15.75 - $19.25/hr
Continuously monitor the health, availability, and utilization of HPC clusters using system dashboards and workload management software (such as Slurm). * Facility Oversight: Monitor critical data ...
Junior Computer Operator
Dayton, OH · On-site
$15.75 - $19.25/hr
Continuously monitor the health, availability, and utilization of HPC clusters using system dashboards and workload management software (such as Slurm). * Facility Oversight: Monitor critical data ...
Junior Computer Operator
Dayton, OH · On-site
$15.75 - $19.25/hr
Continuously monitor the health, availability, and utilization of HPC clusters using system dashboards and workload management software (such as Slurm). * Facility Oversight: Monitor critical data ...
Junior Computer Operator
Dayton, OH · On-site
$15.75 - $19.25/hr
Continuously monitor the health, availability, and utilization of HPC clusters using system dashboards and workload management software (such as Slurm). * Facility Oversight: Monitor critical data ...
AI Infrastructure Operations Engineer
Cleveland, OH · On-site
$104K - $136K/yr
... Slurm, Run:ai * Minimum 6 months hands‑on experience with Claude Code, AI automation tools, Terraform, Ansible, Python, and Bash scripting. * Bachelor's degree or equivalent (minimum 12 years) work ...
AI Infrastructure Operations Engineer
Cleveland, OH · On-site
$104K - $136K/yr
... Slurm, Run:ai * Minimum 6 months hands‑on experience with Claude Code, AI automation tools, Terraform, Ansible, Python, and Bash scripting. * Bachelor's degree or equivalent (minimum 12 years) work ...
Software Engineer I, MLOps (Python/Big Data)
Cleveland, OH · On-site
$111K - $133K/yr
Experience with Kafka, Spark Streaming, or SLURM. * Experience with model lifecycle management. * Experience with cloud-based MLOps platforms, including AWS SageMaker. * Experience in financial ...
Software Engineer I, MLOps (Python/Big Data)
Cleveland, OH · On-site
$111K - $133K/yr
Experience with Kafka, Spark Streaming, or SLURM. * Experience with model lifecycle management. * Experience with cloud-based MLOps platforms, including AWS SageMaker. * Experience in financial ...
Lead integration ofcompute, storage, networking, the AI software stack (CUDA,ROCm, Triton, NIM, NVIDIA AI Enterprise,Run:ai,Slurm, Kubernetes / Kubeflow) and managed-service operating models across ...
Lead integration ofcompute, storage, networking, the AI software stack (CUDA,ROCm, Triton, NIM, NVIDIA AI Enterprise,Run:ai,Slurm, Kubernetes / Kubeflow) and managed-service operating models across ...
Preferred: * 2+ years implementing AI/HPC cluster scheduling (Slurm and Kubernetes), including multi-tenant queues, quotas, and GPU-aware policies * 2+ years supporting generative AI infrastructure ...
Preferred: * 2+ years implementing AI/HPC cluster scheduling (Slurm and Kubernetes), including multi-tenant queues, quotas, and GPU-aware policies * 2+ years supporting generative AI infrastructure ...
Site Rel Engineer Sr. - Model Integration Platform
Cleveland, OH · On-site
$86K - $158K/yr
Linux (RHEL) Administration IBM GPFS / Spectrum Scale Slurm Scheduler Administration Shell/Bash Scripting Basic Python Linux Networking & Security Storage/File Systems (LVM, NAS) Spark / Jupyter ...
Site Rel Engineer Sr. - Model Integration Platform
Cleveland, OH · On-site
$86K - $158K/yr
Linux (RHEL) Administration IBM GPFS / Spectrum Scale Slurm Scheduler Administration Shell/Bash Scripting Basic Python Linux Networking & Security Storage/File Systems (LVM, NAS) Spark / Jupyter ...
Site Rel Engineer Sr. - Model Integration Platform
Cleveland, OH · On-site
$86K - $158K/yr
Linux (RHEL) Administration IBM GPFS / Spectrum Scale Slurm Scheduler Administration Shell/Bash Scripting Basic Python Linux Networking & Security Storage/File Systems (LVM, NAS) Spark / Jupyter ...
Site Rel Engineer Sr. - Model Integration Platform
Cleveland, OH · On-site
$86K - $158K/yr
Linux (RHEL) Administration IBM GPFS / Spectrum Scale Slurm Scheduler Administration Shell/Bash Scripting Basic Python Linux Networking & Security Storage/File Systems (LVM, NAS) Spark / Jupyter ...
Preferred: * 2+ years implementing AI/HPC cluster scheduling (Slurm and Kubernetes), including multi-tenant queues, quotas, and GPU-aware policies * 2+ years supporting generative AI infrastructure ...
Preferred: * 2+ years implementing AI/HPC cluster scheduling (Slurm and Kubernetes), including multi-tenant queues, quotas, and GPU-aware policies * 2+ years supporting generative AI infrastructure ...
Slurm information
What are popular job titles related to Slurm jobs in Ohio?
For Slurm jobs in Ohio, the most frequently searched job titles are:
What job categories do people searching Slurm jobs in Ohio look for?
The top searched job categories for Slurm jobs in Ohio are:

Machine Learning Engineer
Bowling Green, OH • On-site
Full-time
Re-posted 26 days ago
Job description
-
Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch
-
Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration
-
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues
-
Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy
-
Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams
-
Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics
-
Implement methods from recent ML papers quickly and turn them into production-grade systems