Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Gary, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Gary, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Evansville, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Evansville, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
South Bend, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
South Bend, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Terre Haute, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Terre Haute, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Hammond, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Hammond, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Fort Wayne, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
Machine Learning Engineer
Fort Wayne, IN · On-site
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...
HPC Engineer
Indianapolis, IN · On-site
$75K - $80K/yr
... Slurm job scheduler. • Excellent troubleshooting skills with the ability to resolve application-related issues. • Strong documentation and diagramming abilities. • Ability to work ...
HPC Engineer
Indianapolis, IN · On-site
$75K - $80K/yr
... Slurm job scheduler. • Excellent troubleshooting skills with the ability to resolve application-related issues. • Strong documentation and diagramming abilities. • Ability to work ...
AI PhD ML Engineering Intern
Indianapolis, IN · On-site
$16 - $20.75/hr
Exposure to batch/distributed job execution in at least one of: traditional HPC job schedulers (e.g., Slurm, Grid Engine) or running batch workloads in a Kubernetes environment * Available for 12 ...
AI PhD ML Engineering Intern
Indianapolis, IN · On-site
$16 - $20.75/hr
Exposure to batch/distributed job execution in at least one of: traditional HPC job schedulers (e.g., Slurm, Grid Engine) or running batch workloads in a Kubernetes environment * Available for 12 ...
AI PhD ML Engineering Intern
Indianapolis, IN · On-site
$16 - $20.75/hr
Exposure to batch/distributed job execution in at least one of: traditional HPC job schedulers (e.g., Slurm, Grid Engine) or running batch workloads in a Kubernetes environment * Available for 12 ...
AI PhD ML Engineering Intern
Indianapolis, IN · On-site
$16 - $20.75/hr
Exposure to batch/distributed job execution in at least one of: traditional HPC job schedulers (e.g., Slurm, Grid Engine) or running batch workloads in a Kubernetes environment * Available for 12 ...
AI PhD ML Engineering Intern
Indianapolis, IN · On-site
$16 - $20.75/hr
Exposure to batch/distributed job execution in at least one of: traditional HPC job schedulers (e.g., Slurm, Grid Engine) or running batch workloads in a Kubernetes environment * Available for 12 ...
AI PhD ML Engineering Intern
Indianapolis, IN · On-site
$16 - $20.75/hr
Exposure to batch/distributed job execution in at least one of: traditional HPC job schedulers (e.g., Slurm, Grid Engine) or running batch workloads in a Kubernetes environment * Available for 12 ...
Slurm information
What are popular job titles related to Slurm jobs in Indiana?
For Slurm jobs in Indiana, the most frequently searched job titles are:
What job categories do people searching Slurm jobs in Indiana look for?
The top searched job categories for Slurm jobs in Indiana are:

Machine Learning Engineer
Gary, IN
Full-time
Re-posted 28 days ago
Job description
-
Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch
-
Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration
-
Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues
-
Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy
-
Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams
-
Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics
-
Implement methods from recent ML papers quickly and turn them into production-grade systems