1

Machine Learning Infrastructure Jobs (NOW HIRING)

next page

Showing results 1-20

Machine Learning Infrastructure information

See salary details

$15

$28

$52

How much do machine learning infrastructure jobs pay per hour?

As of Sep 12, 2026, the average hourly pay for machine learning infrastructure in the United States is $28.01, according to ZipRecruiter salary data. Most workers in this role earn between $21.88 and $30.29 per hour, depending on experience, location, and employer.

What are the typical challenges faced by professionals working in machine learning infrastructure roles?

Professionals in Machine Learning Infrastructure often encounter challenges related to scaling systems to handle large datasets, ensuring model reproducibility, and maintaining efficient workflows for both development and deployment. Collaborating closely with data scientists, software engineers, and DevOps teams is crucial to address issues like version control, resource allocation, and performance optimization. Staying updated on evolving tools and cloud platforms is also essential, as the landscape changes rapidly and impacts system design and integration.

What are the key skills and qualifications needed to thrive in machine learning infrastructure, and why are they important?

To excel in Machine Learning Infrastructure, you need a solid background in computer science, software engineering, and distributed systems, often supported by experience in deploying and scaling machine learning models. Familiarity with cloud platforms (like AWS, GCP, or Azure), containerization tools (such as Docker and Kubernetes), and ML workflow systems (e.g., TensorFlow Extended, MLflow) is crucial. Strong problem-solving skills, collaboration, and the ability to communicate technical concepts effectively help you stand out in this field. These skills ensure scalable, reliable, and efficient deployment of ML solutions, enabling organizations to leverage machine learning at production scale.

What is the difference between Machine Learning Infrastructure vs Data Engineer?

AspectMachine Learning InfrastructureData Engineer
Required CredentialsBachelor's in CS, experience with ML toolsBachelor's in CS, experience with data pipelines
Work EnvironmentFocus on ML systems, cloud platformsData pipelines, database management
Employer & Industry UsageTech companies, AI startupsAny industry with data needs, tech firms
Search & Comparison IntentUnderstanding ML system setupBuilding data pipelines

Machine Learning Infrastructure specialists focus on deploying and maintaining systems that support machine learning models, often working with cloud platforms and ML tools. Data Engineers build and manage data pipelines and databases, supporting data collection and processing. While both roles require technical skills and overlap in data handling, Machine Learning Infrastructure is more centered on ML system deployment, whereas Data Engineers focus on data architecture and pipelines.

What does a machine learning infrastructure engineer do?

A machine learning infrastructure engineer designs, builds, and maintains the systems and tools that support machine learning workflows, including data pipelines, model deployment, and scalable computing resources. They often work with cloud platforms, containerization, and automation tools to ensure efficient and reliable model training and deployment environments.
More about Machine Learning Infrastructure jobs
Infographic showing various Machine Learning Infrastructure job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 1% As Needed, 74% Full Time, 22% Part Time, and 2% Contract. Highlights an 83% Physical, 2% Hybrid, and 15% Remote job distribution, with an average salary of $58,269 per year, or $28 per hour.

Machine Learning Infrastructure Engineer

Sunnyvale, CA • On-site

MBZUAI (Mohamed bin Zayed University of Artificial Intelligence)

$125K - $164K/yr

Full-time

Re-posted 12 days ago


Job description

Job Summary:
MBZUAI is a dedicated research lab focused on advancing AI research and building high-performance computing systems. The Machine Learning Infrastructure Engineer will work on extending and scaling training systems, collaborating with researchers and engineers to tackle challenges in AI development.
Responsibilities:
• Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
• Implement distributed optimizers from mathematical specs
• Build robust config + launch systems across multi-node, multi-GPU clusters
• Own experiment tracking, metrics logging, and job monitoring for external visibility
• Improve training system reliability, maintainability, and performance
• Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures
• Translate mathematical optimizer specs into distributed implementations
• Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets
• Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers
• Write production-quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.
Qualifications:
Required:
• 5+ years of experience in ML systems, infra, or distributed training
• Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
• Strong software engineering fundamentals (Python, systems design, testing)
• Proven multi-node experience (e.g., Slurm, Kubernetes, Ray) and debugging skills (e.g., NCCL/GLOO)
• Ability to implement algorithms across GPUs/nodes based on mathematical specs
• Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team
• Experience with large-scale machine learning workloads (strong ML fundamentals)
Preferred:
• Exposure to mixed-precision training (e.g., bf16, fp8) with accuracy validation
• Familiarity with performance profiling, kernel fusion, or memory optimization
• Open-source contributions or published research (MLSys, ICML, NeurIPS)
• CUDA or Triton kernel experience
• Experience with large-scale pre-training
• Experience building custom training pipelines at scale and modifying them for custom needs
• Deep familiarity with training infrastructure and performance tuning
Company:
Official account of Mohamed bin Zayed University of Artificial Intelligence. Dedicated to research, innovation, and empowering brilliant minds in AI. Founded in 2019, the company is headquartered in Abu Dhabi, ARE, with a team of 51-200 employees. The company is currently Growth Stage.