1

Slurm Jobs in Florida (NOW HIRING)

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues * Build evaluation harnesses and benchmark infrastructure, with held ...

Showing results 41-60

Machine Learning Engineer

Bespoke Labs

Hialeah, FL • On-site

Full-time

Re-posted 22 days ago


Job description

  • Own model training and post-training pipelines end to end: SFT, RLHF, PPO, DPO, and reward model training in PyTorch

  • Build and maintain the infrastructure around RL training: rollout collection, data curation, reward model serving, and experiment orchestration

  • Run and scale training experiments on cloud or HPC (AWS, GCP, SLURM, Ray), and debug throughput, stability, and convergence issues

  • Build evaluation harnesses and benchmark infrastructure, with held-out sets and contamination controls, so results are trustworthy

  • Read eval signal and training curves to determine whether a change actually helped, and feed findings back to the research and environment teams

  • Integrate RL environments into the training stack, working with environment authors on interfaces, reward plumbing, and agent loop mechanics

  • Implement methods from recent ML papers quickly and turn them into production-grade systems