1

Slurm Jobs in Ohio (NOW HIRING)

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Setting up, configuring, and continuously evolving NEURA's HyperPod clusters, including HyperPod/Slurm and HyperPod/EKS orchestration models. * Designing and implementing strategies for cluster ...

Lead integration ofcompute, storage, networking, the AI software stack (CUDA,ROCm, Triton, NIM, NVIDIA AI Enterprise,Run:ai,Slurm, Kubernetes / Kubeflow) and managed-service operating models across ...

Junior Computer Operator

Dayton, OH · On-site

$15.75 - $19.25/hr

Continuously monitor the health, availability, and utilization of HPC clusters using system dashboards and workload management software (such as Slurm). * Facility Oversight: Monitor critical data ...

Junior Computer Operator

Dayton, OH · On-site

$15.75 - $19.25/hr

Continuously monitor the health, availability, and utilization of HPC clusters using system dashboards and workload management software (such as Slurm). * Facility Oversight: Monitor critical data ...

Junior Computer Operator

Dayton, OH · On-site

$15.75 - $19.25/hr

Continuously monitor the health, availability, and utilization of HPC clusters using system dashboards and workload management software (such as Slurm). * Facility Oversight: Monitor critical data ...

AI Infrastructure Operations Engineer

Cleveland, OH · On-site

$104K - $136K/yr

... Slurm, Run:ai * Minimum 6 months hands‑on experience with Claude Code, AI automation tools, Terraform, Ansible, Python, and Bash scripting. * Bachelor's degree or equivalent (minimum 12 years) work ...

Lead integration ofcompute, storage, networking, the AI software stack (CUDA,ROCm, Triton, NIM, NVIDIA AI Enterprise,Run:ai,Slurm, Kubernetes / Kubeflow) and managed-service operating models across ...

Preferred: * 2+ years implementing AI/HPC cluster scheduling (Slurm and Kubernetes), including multi-tenant queues, quotas, and GPU-aware policies * 2+ years supporting generative AI infrastructure ...

Preferred: * 2+ years implementing AI/HPC cluster scheduling (Slurm and Kubernetes), including multi-tenant queues, quotas, and GPU-aware policies * 2+ years supporting generative AI infrastructure ...

next page

Showing results 1-20

Slurm information

What are popular job titles related to Slurm jobs in Ohio?

For Slurm jobs in Ohio, the most frequently searched job titles are:

Infographic showing various Slurm job openings in Ohio as of August 2026, with employment types broken down into 90% Full Time, 6% Part Time, and 4% Contract. Highlights an 79% Physical, 7% Hybrid, and 14% Remote job distribution.

Machine Learning Engineer

Bowling Green, OH • On-site

Full-time

Re-posted 26 days ago


Job description

  • Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch

  • Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration

  • Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues

  • Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy

  • Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams

  • Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics

  • Implement methods from recent ML papers quickly and turn them into production-grade systems