1

Ml Infrastructure Jobs (NOW HIRING)

Senior ML Infrastructure Engineer

New York, NY ยท On-site

$180K - $230K/yr

We're hiring a Senior ML Infrastructure Engineer to build and own the infrastructure that powers it - from training models on tens of millions of patients and hundreds of millions of rows of claims ...

Software Engineer - ML Infrastructure

San Francisco, CA ยท On-site

$203K - $241K/yr

The Role Specter is hiring an ML Infrastructure engineer to build and scale the machine learning systems that power real-time perception and inference across our edge-cloud platform. This role owns ...

ML Infrastructure Engineer

Sunnyvale, CA ยท Hybrid

$119K - $187K/yr

Hands-on experience in ML platforms * Experience with GPU/TPU optimizations * Experience with Ray framework * Experience with Kubernetes at Scale * Experience infrastructure applications or similar ...

$119 - $188/hr

Hands-on experience in ML platforms * Experience with GPU/TPU optimizations * Experience with Ray framework * Experience with Kubernetes at Scale * Experience infrastructure applications or similar ...

New

ML Infrastructure Engineer

Sunnyvale, CA ยท On-site

$119 - $188/hr

Hands-on experience in ML platforms * Experience with GPU/TPU optimizations * Experience with Ray framework * Experience with Kubernetes at Scale * Experience infrastructure applications or similar ...

ML Infrastructure Engineer

Sunnyvale, CA ยท On-site

$119K - $187K/yr

Hands-on experience in ML platforms * Experience with GPU/TPU optimizations * Experience with Ray framework * Experience with Kubernetes at Scale * Experience infrastructure applications or similar ...

Showing results 41-60

Ml Infrastructure information

See salary details

$46.5K

$127.1K

$182K

How much do ml infrastructure jobs pay per year?

As of Aug 23, 2026, the average yearly pay for ml infrastructure in the United States is $127,066.00, according to ZipRecruiter salary data. Most workers in this role earn between $107,500.00 and $141,000.00 per year, depending on experience, location, and employer.

What is ML infrastructure?

ML Infrastructure refers to the underlying systems, tools, and processes that enable the development, deployment, and scaling of machine learning models. This includes data storage and management, computing resources, model training and serving environments, monitoring, and automation tools. ML Infrastructure ensures that data scientists and engineers can efficiently build, test, and maintain machine learning applications in a reliable and reproducible manner. It is a crucial foundation for organizations looking to operationalize AI and machine learning solutions at scale.

What are some common challenges faced by professionals working in ML infrastructure roles?

Professionals in ML Infrastructure often encounter challenges related to scaling systems to handle large volumes of data, ensuring reliable deployment pipelines, and maintaining reproducibility across different environments. They must also collaborate closely with data scientists and engineers to streamline workflows and address issues like version control and model monitoring. Staying updated with rapidly evolving tools and best practices is essential, and balancing stability with innovation is a frequent aspect of the role.

What are the key skills and qualifications needed to thrive as an ML infrastructure engineer, and why are they important?

To thrive as an ML Infrastructure Engineer, you need a strong background in software engineering, cloud computing, and machine learning concepts, often supported by a degree in computer science or a related field. Proficiency with containerization tools (like Docker and Kubernetes), cloud platforms (such as AWS, GCP, or Azure), and CI/CD systems is critical. Excellent problem-solving, collaboration, and communication skills help you efficiently work with data scientists and DevOps teams. These skills and qualities are vital for building scalable, reliable ML systems that support rapid experimentation and deployment in production environments.

What is the difference between Ml Infrastructure vs Data Engineer?

AspectML InfrastructureData Engineer
Required CredentialsBachelor's in CS, Data Science, or related; knowledge of cloud platformsBachelor's in CS, Software Engineering, or related; experience with databases and ETL tools
Work EnvironmentFocus on deploying and maintaining ML systems, cloud environments, and infrastructure toolsDesigning, building, and managing data pipelines and storage solutions
Industry UsageUsed in AI/ML teams to support model deployment and scalabilityUsed across data-driven organizations for data management and analytics

ML Infrastructure specialists focus on deploying, scaling, and maintaining machine learning systems and infrastructure, while Data Engineers primarily build and manage data pipelines and storage solutions. Both roles require technical skills and often collaborate, but their core responsibilities differ in focus and tools used.

More about Ml Infrastructure jobs

What cities are hiring for Ml Infrastructure jobs?

Cities with the most Ml Infrastructure job openings:

What states have the most Ml Infrastructure jobs?

States with the most job openings for Ml Infrastructure jobs include:

Infographic showing various Ml Infrastructure job openings in the United States as of August 2026, with employment types broken down into 91% Full Time, 5% Part Time, and 4% Contract. Highlights an 83% Physical, 6% Hybrid, and 11% Remote job distribution, with an average salary of $127,066 per year, or $61.1 per hour.

ML Infrastructure Engineer, Training

DYNA Robotics Inc

Redwood City, CA โ€ข On-site

$220K - $320K/yr

Full-time

Re-posted 6 days ago


Job description

Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model with top-in-industry generalization and real-world performance. Already deployed with customers across multiple industries, our robots do commercial-grade work in the physical world. Our team comes from Google DeepMind, Meta, and Cruise, and we're backed by CRV, First Round, and other leading investors.
The Role
As a ML Training Infrastructure Engineer, you will architect and build the systems that turn our multi-cloud GPU fleet into a training engine our researchers love. Your charter is singular and broad: own training infrastructure end-to-end so that every GPU is busy, every run is reproducible, and every researcher's next experiment is one command away.
What You'll Do
  • Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters. You'll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.
  • Optimize Researcher Ergonomics: Build a research codebase and job scheduling system (Kubernetes/SLURM) that prioritizes fast iteration, automated retries, and seamless failure recovery.
  • High-Performance Data Handling: Design high-throughput pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals), ensuring dataloaders never starve the GPUs.
  • Production Inference: Build low-latency inference pipelines for real-time robot control. You'll apply quantization, distillation, and model compilation (TensorRT, Triton) to move models from the lab to the physical world.
  • Deep Systems Profiling: Dive into the weeds of GPU utilization, I/O bottlenecks, and memory fragmentation to squeeze every bit of performance out of our expanding compute fleet.
What You'll Bring
  • 7+ Years of Engineering: With a track record of leading technical projects in high-performance computing (HPC) or ML infrastructure.
  • ML Systems Mastery: Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate). You understand the nuances of mixed precision and gradient accumulation.
  • Infrastructure Expertise: Hands-on experience managing cloud GPU environments (GCP/AWS) and container orchestration (Kubernetes).
  • Low-Level Intuition: A fundamental understanding of distributed systems, including race conditions, memory management, and NCCL/inter-node communication.
  • Ownership Mindset: You don't just "deploy" code; you design, build, and operate systems end-to-end to unblock fast-moving research.
Bonus Points For
  • Experience with Robotics Data Formats (MCAP, Protobuf) or multimodal models (VLAs).
  • Deep ML systems experience: custom kernels (Triton), compilers, or runtime optimization.
  • Experience as a founding or early-stage infrastructure hire.

At Dyna Robotics, we build technology for the real world, which requires a team as diverse as the environments our robots inhabit. We are an equal opportunity employer committed to technical rigor and mutual respect.
Don't let a checklist stop you. Data shows that underrepresented groups often only apply if they meet 100% of the criteria. We value problem-solving and grit over keyword matching. If you're passionate about the intersection of geometry and robotics, we want to hear from you-even if you don't check every box.