1

Slurm Jobs in Missouri (NOW HIRING)

Experience with workload managers (SLURM, Argo Workflow, Airflow). * Strong problem-solving and troubleshooting skills. * Strong communication and interpersonal skills. * Must possess excellent time ...

Experience with workload managers (SLURM, Argo Workflow, Airflow). * Strong problem-solving and troubleshooting skills. * Strong communication and interpersonal skills. * Must possess excellent time ...

$48.50 - $64/hr

Familiarity with technologies including Kubernetes, Slurm, Docker, Helm, Git, Terraform, and infrastructure-as-code practices. * Excellent written and verbal communication skills, with the ability to ...

AI and ML Infra Software Engineer, GPU Clusters

Jobtailor

California, MO • On-site

$120 - $190/hr

Other

This job post has expired today. Applications are no longer accepted.


Job description

  • Collaborate closely with our AI and ML research teams to understand their infrastructure needs and obstacles
  • Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization
  • Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results
  • Collaborate with diverse teams, including researchers, data engineers, and DevOps professionals, to build a seamless and coordinated AI/ML infrastructure ecosystem
  • Stay on top of the latest advancements in AI/ML technologies, frameworks, and effective strategies, and promote their implementation within the company
Requirements
  • Recent graduate with a MS, PhD or equivalent experience in Computer Science or related field
  • Proven experience in AI/ML and HPC workloads and infrastructure
  • Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure
  • In-depth knowledge of accelerated computing (e.g., GPU, custom silicon)
  • Storage (e.g., Lustre, GPFS, BeeGFS)
  • Scheduling & orchestration (e.g., Slurm, Kubernetes, LSF)
  • High-speed networking (e.g., Infiniband, RoCE, Amazon EFA)
  • Containers technologies (Docker, Enroot)
  • Expertise in running and optimizing large-scale distributed training workloads using PyTorch (DDP, FSDP), NeMo, or JAX
  • Deep understanding of AI/ML workflows, encompassing data processing, model training, and inference pipelines
  • Proficiency in programming & scripting languages such as Python, Go, Bash
  • Familiarity with cloud computing platforms (e.g., AWS, GCP, Azure)
  • Experience with parallel computing frameworks and paradigms.
  • Passion for continual learning and keeping abreast of new technologies and effective approaches in the AI/ML infrastructure field.
  • Excellent communication and collaboration skills
Core Competencies

Demonstrates expertise in AI/ML infrastructure, including High Performance Computing (HPC) and accelerated computing technologies. Proficient in optimizing large-scale distributed training workloads and collaborating with diverse teams to enhance AI researcher efficiency.

Highest-signal resume keywords
  • AI/ML Infrastructure
  • High Performance Computing (HPC)
  • Distributed Training Workloads Optimization
  • Programming Languages (Python, Go, Bash)
  • Cloud Computing Platforms (AWS, GCP, Azure)
ATS Optimization KeywordsHard Skills
  • AI/ML Workloads
  • Accelerated Computing
  • Storage Technologies (Lustre, GPFS, BeeGFS)
  • Scheduling & Orchestration (Slurm, Kubernetes, LSF)
  • High-Speed Networking (Infiniband, RoCE, Amazon EFA)
  • Container Technologies (Docker, Enroot)
  • Parallel Computing Frameworks
  • Data Processing
  • Model Training
  • Inference Pipelines
Soft Skills
  • Excellent Communication
  • Collaboration Skills
  • Passion for Learning
Certifications & Qualifications
  • MS or PhD in Computer Science or Related Field
#J-18808-Ljbffr