1

Machine Learning Infrastructure Engineer Jobs (NOW HIRING)

next page

Showing results 1-20

Machine Learning Infrastructure Engineer information

See salary details

$46.5K

$127.1K

$182K

How much do machine learning infrastructure engineer jobs pay per year?

As of Aug 22, 2026, the average yearly pay for machine learning infrastructure engineer in the United States is $127,066.00, according to ZipRecruiter salary data. Most workers in this role earn between $107,500.00 and $141,000.00 per year, depending on experience, location, and employer.

What is a machine learning infrastructure engineer?

A Machine Learning Infrastructure Engineer designs, builds, and maintains the systems that support the development and deployment of machine learning models. This includes managing data pipelines, optimizing model training and inference, and ensuring scalability and reliability in production environments. They work closely with data scientists, ML engineers, and DevOps teams to create efficient workflows and infrastructure. Key technologies often include cloud platforms, containerization, orchestration tools, and distributed computing frameworks.

What are the key skills and qualifications needed to thrive as a machine learning infrastructure engineer?

To thrive as a Machine Learning Infrastructure Engineer, you need a strong background in computer science, cloud computing, distributed systems, and experience with machine learning frameworks, often supported by a degree in a related field. Familiarity with tools such as Docker, Kubernetes, Terraform, as well as cloud platforms like AWS, GCP, or Azure, and certifications in cloud or DevOps technologies are highly valued. Strong problem-solving abilities, effective communication, and collaboration skills help engineers work seamlessly with data scientists and cross-functional teams. These skills are essential to design, implement, and maintain robust, scalable infrastructure that enables efficient machine learning development and deployment.

What are some common challenges faced by machine learning infrastructure engineers, and how can these be addressed on the job?

Machine Learning Infrastructure Engineers often face challenges such as ensuring infrastructure scalability, managing resource allocation, and maintaining system reliability while supporting rapid experimentation by data science teams. Balancing the needs for flexibility in research environments with production-grade stability requires a deep understanding of both engineering best practices and the unique requirements of machine learning workflows. Collaboration with data scientists, clear communication about infrastructure capabilities, and staying current with fast-evolving technologies are key strategies for success. Most companies encourage ongoing learning and provide opportunities to contribute to architecture decisions, which makes this a rewarding environment for problem-solvers and innovators.

More about Machine Learning Infrastructure Engineer jobs

What cities are hiring for Machine Learning Infrastructure Engineer jobs?

Cities with the most Machine Learning Infrastructure Engineer job openings:

What states have the most Machine Learning Infrastructure Engineer jobs?

States with the most job openings for Machine Learning Infrastructure Engineer jobs include:

Infographic showing various Machine Learning Infrastructure Engineer job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 75% Full Time, 23% Part Time, and 1% Contract. Highlights an 87% Physical, 2% Hybrid, and 11% Remote job distribution, with an average salary of $127,066 per year, or $61.1 per hour.

Staff Machine Learning Infrastructure Engineer

Atoms

San Francisco, CA • On-site

$126K - $166K/yr

Full-time

Re-posted 21 days ago


Job description

Job Summary:
Atoms is building the machines that power the next era of progress. They are seeking a foundational Machine Learning Infrastructure Engineer to design and build the large-scale ML training infrastructure that powers their next-generation autonomous transport models.
Responsibilities:
• Design, implement, and scale repeatable machine learning infrastructure utilizing Kubernetes to support large-scale distributed GPU training of novel neural networks.
• Leverage distributed compute frameworks to efficiently manage and execute a high volume of complex ML training jobs concurrently across large GPU clusters.
• Integrate advanced model management and experiment tracking tools to provide researchers with deep observability into training metrics and run performance.
• Build and optimize high-throughput data ingestion pipelines to seamlessly stream petabyte-scale multi-sensor vehicle logs into training environments.
• Architect robust infrastructure for autonomous model validation and continuous integration testing, ensuring new vehicle policy releases are entirely regression-free.
• Partner closely with core robotics engineers and machine learning researchers to eliminate workflow bottlenecks and accelerate the deploy-to-vehicle lifecycle.
Qualifications:
Required:
• 8+ years of professional software engineering career experience
• Strong backend systems programming skills with proficiency in Go, Python, Java or similar
• Proficiency with Kubernetes for container orchestration and building cloud-agnostic environments from scratch
• Experience implementing distributed ML compute frameworks (e.g., Ray) to coordinate large pools of GPUs for heavy, multi-node workloads
• Hands-on experience building MLOps pipelines, metadata tracking architectures, and model registries using platforms like MLflow
• Prior experience managing high-throughput data pipelines using modern distributed data engines to feed data-hungry neural network architectures
Preferred:
• Familiarity or exposure to Rust considered a plus.
Company:
Atoms is a robotics startup that develops industrial robotics and physical AI systems to automate tasks across various industries. Founded in 2016, the company is headquartered in Los Angeles, USA, with a team of 1001-5000 employees. The company is currently Late Stage.