1

Machine Learning Infrastructure Engineer Jobs (NOW HIRING)

AI Infrastructure Engineer MLOps

San Francisco, CA ยท On-site

$126K - $166K/yr

... machine learning platforms running in cloud native environments. Responsibilities * Design, deploy, and support scalable AI/ML infrastructure platforms * Manage Kubernetes environments running in AWS ...

Showing results 21-40

Machine Learning Infrastructure Engineer information

See salary details

$46.5K

$127.1K

$182K

How much do machine learning infrastructure engineer jobs pay per year?

As of Sep 12, 2026, the average yearly pay for machine learning infrastructure engineer in the United States is $127,066.00, according to ZipRecruiter salary data. Most workers in this role earn between $107,500.00 and $141,000.00 per year, depending on experience, location, and employer.

What is a machine learning infrastructure engineer?

A Machine Learning Infrastructure Engineer designs, builds, and maintains the systems that support the development and deployment of machine learning models. This includes managing data pipelines, optimizing model training and inference, and ensuring scalability and reliability in production environments. They work closely with data scientists, ML engineers, and DevOps teams to create efficient workflows and infrastructure. Key technologies often include cloud platforms, containerization, orchestration tools, and distributed computing frameworks.

What are the key skills and qualifications needed to thrive as a machine learning infrastructure engineer?

To thrive as a Machine Learning Infrastructure Engineer, you need a strong background in computer science, cloud computing, distributed systems, and experience with machine learning frameworks, often supported by a degree in a related field. Familiarity with tools such as Docker, Kubernetes, Terraform, as well as cloud platforms like AWS, GCP, or Azure, and certifications in cloud or DevOps technologies are highly valued. Strong problem-solving abilities, effective communication, and collaboration skills help engineers work seamlessly with data scientists and cross-functional teams. These skills are essential to design, implement, and maintain robust, scalable infrastructure that enables efficient machine learning development and deployment.

What are some common challenges faced by machine learning infrastructure engineers, and how can these be addressed on the job?

Machine Learning Infrastructure Engineers often face challenges such as ensuring infrastructure scalability, managing resource allocation, and maintaining system reliability while supporting rapid experimentation by data science teams. Balancing the needs for flexibility in research environments with production-grade stability requires a deep understanding of both engineering best practices and the unique requirements of machine learning workflows. Collaboration with data scientists, clear communication about infrastructure capabilities, and staying current with fast-evolving technologies are key strategies for success. Most companies encourage ongoing learning and provide opportunities to contribute to architecture decisions, which makes this a rewarding environment for problem-solvers and innovators.

More about Machine Learning Infrastructure Engineer jobs

What cities are hiring for Machine Learning Infrastructure Engineer jobs?

Cities with the most Machine Learning Infrastructure Engineer job openings:

What states have the most Machine Learning Infrastructure Engineer jobs?

States with the most job openings for Machine Learning Infrastructure Engineer jobs include:

What are popular job titles related to Machine Learning Infrastructure Engineer jobs?

For Machine Learning Infrastructure Engineer jobs, the most frequently searched job titles are:

Infographic showing various Machine Learning Infrastructure Engineer job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 1% As Needed, 74% Full Time, 22% Part Time, and 2% Contract. Highlights an 83% Physical, 2% Hybrid, and 15% Remote job distribution, with an average salary of $127,066 per year, or $61.1 per hour.

Staff Machine Learning Infrastructure Engineer

San Francisco, CA โ€ข On-site

Atoms
Specialized Design Servicesย โ€ขย 11 - 50 employees

$224K - $280K/yr

Full-time

Medical, Dental, Vision, Life, Retirement

Re-posted 12 days ago


Job description

Who we areย 

Atoms is building the machines that power the next era of progress.

Over the last decade, software has transformed the digital world. But the physical world, where food is made, minerals are mined, goods are moved, and industries are run, remains far less intelligent, far less efficient, and far more constrained. We're changing that.

Atoms builds Physical AI- real-world robots for the industries that move civilization forward, starting with food, mining, and transport. Our systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive.

This work requires more than robotics. It requires deep integration across hardware, software, AI, operations, manufacturing, and real estate. We don't just build machines in a lab. We deploy them into real environments, operate them, learn from them, and improve them until they work at scale.

We are roboticists, engineers, operators, and builders. We believe the next great technology companies will not only transform information, but the physical systems that shape everyday life.

If you want to work on hard problems with real-world impact, join us.


What you'll do

We are seeking a foundational Staff Machine Learning Infrastructure Engineer to design and build the large-scale ML training infrastructure that powers our next-generation autonomous transport models. In this role, you will design the high-performance training pipelines and validation environments that enable our world-class robotics and ML researchers to iterate rapidly. You will own the challenge of scaling distributed GPU workloads to support a high volume of concurrent training runs across an expanding vehicle fleet, building a platform that can flexibly run on whatever GPU capacity is available, regardless of provider or environment, directly accelerating innovation across the platform.

  • Training Infrastructure: Design, implement, and scale repeatable machine learning infrastructure utilizing Kubernetes to support large-scale distributed GPU training of novel neural networks.
  • Distributed Computing & Orchestration: Leverage distributed compute frameworks to efficiently manage and execute a high volume of complex ML training jobs concurrently across large GPU clusters.
  • Experiment Tracking & MLOps: Integrate advanced model management and experiment tracking tools to provide researchers with deep observability into training metrics and run performance.
  • Data Engineering Pipelines: Build and optimize high-throughput data ingestion pipelines to seamlessly stream petabyte-scale multi-sensor vehicle logs into training environments.
  • Validation at Scale: Architect robust infrastructure for autonomous model validation and continuous integration testing, ensuring new vehicle policy releases are entirely regression-free.
  • Cross-Functional Collaboration: Partner closely with core robotics engineers and machine learning researchers to eliminate workflow bottlenecks and accelerate the deploy-to-vehicle lifecycle.
ย 

What we're looking for

  • 8+ years of professional software engineering career experience
  • Strong backend systems programming skills with proficiency in Go, Python, Java or similar (with familiarity or exposure to Rust considered a plus).
  • Proficiency with Kubernetes for container orchestration and building cloud-agnostic environments from scratch.
  • Experience implementing distributed ML compute frameworks (e.g., Ray) to coordinate large pools of GPUs for heavy, multi-node workloads.
  • Hands-on experience building MLOps pipelines, metadata tracking architectures, and model registries using platforms like MLflow.
  • Prior experience managing high-throughput data pipelines using modern distributed data engines to feed data-hungry neural network architectures.

Why join us

At Atoms, you'll work on one of the defining challenges of our time - bringing automation into the physical world to drive real, lasting impact. We exist to uncover valuable unknown truths and turn them into progress, which means constantly pushing beyond what's known and building what doesn't yet exist. The work is ambitious and often challenging, but it's grounded in a shared sense of purpose and a team committed to seeing it through together. Our work only matters if it serves others, and we know that meaningful progress depends on the trust of the people we serve and the strength of our team-so we invest in both, creating an environment where you can do your best work and grow.

What else you need to know

This role is based in our San Francisco office. Atoms is a company driven by invention and continuous change - we are constantly reimagining our industries, building new products, and refining how we operate. We do our best work together. That's why all of our office-based teams work onsite, five days a week.

The base salary range for this role is $224,000 - $280,000 per year.ย 

Actual compensation will be determined on an individual basis and may vary depending on experience, skills, and qualifications.

Base salary is just one part of your total rewards package. You may also be eligible for equity awards.

Benefits Summary (USA Full-Time Exempt Employees):

  • Medical, Dental, Vision, Disability, and Life Insurance
  • Flexible Spending Account / Health Savings Account Options
  • 401(k)
  • Equity
  • Sick Time, Unlimited Flexible Time Off, and Paid Holidays
  • Paid Parental Leaveย 
  • Pre-Tax Commuter Benefit Plan
  • Team lunch in our SoMa office every Tuesday and Thursday

Benefits are subject to change at the company's discretion.
Atoms accepts applications on an ongoing basis.

Ready to join us as we serve those who serve others?ย 

#LI-Onsite