1

Machine Learning Infrastructure Engineer Jobs in San Jose, CA

next page

Showing results 1-20

Machine Learning Infrastructure Engineer information

See San Jose, CA salary details

$54.5K

$148.9K

$213.3K

How much do machine learning infrastructure engineer jobs pay per year?

As of Aug 22, 2026, the average yearly pay for machine learning infrastructure engineer in San Jose, CA is $148,920.00, according to ZipRecruiter salary data. Most workers in this role earn between $126,000.00 and $165,200.00 per year, depending on experience, location, and employer.

What is a machine learning infrastructure engineer?

A Machine Learning Infrastructure Engineer designs, builds, and maintains the systems that support the development and deployment of machine learning models. This includes managing data pipelines, optimizing model training and inference, and ensuring scalability and reliability in production environments. They work closely with data scientists, ML engineers, and DevOps teams to create efficient workflows and infrastructure. Key technologies often include cloud platforms, containerization, orchestration tools, and distributed computing frameworks.

What are the key skills and qualifications needed to thrive as a machine learning infrastructure engineer?

To thrive as a Machine Learning Infrastructure Engineer, you need a strong background in computer science, cloud computing, distributed systems, and experience with machine learning frameworks, often supported by a degree in a related field. Familiarity with tools such as Docker, Kubernetes, Terraform, as well as cloud platforms like AWS, GCP, or Azure, and certifications in cloud or DevOps technologies are highly valued. Strong problem-solving abilities, effective communication, and collaboration skills help engineers work seamlessly with data scientists and cross-functional teams. These skills are essential to design, implement, and maintain robust, scalable infrastructure that enables efficient machine learning development and deployment.

What are some common challenges faced by machine learning infrastructure engineers, and how can these be addressed on the job?

Machine Learning Infrastructure Engineers often face challenges such as ensuring infrastructure scalability, managing resource allocation, and maintaining system reliability while supporting rapid experimentation by data science teams. Balancing the needs for flexibility in research environments with production-grade stability requires a deep understanding of both engineering best practices and the unique requirements of machine learning workflows. Collaboration with data scientists, clear communication about infrastructure capabilities, and staying current with fast-evolving technologies are key strategies for success. Most companies encourage ongoing learning and provide opportunities to contribute to architecture decisions, which makes this a rewarding environment for problem-solvers and innovators.

What are popular job titles related to Machine Learning Infrastructure Engineer jobs in San Jose, CA?

For Machine Learning Infrastructure Engineer jobs in San Jose, CA, the most frequently searched job titles are:

What job categories do people searching Machine Learning Infrastructure Engineer jobs in San Jose, CA look for?

The top searched job categories for Machine Learning Infrastructure Engineer jobs in San Jose, CA are:

What cities near San Jose, CA are hiring for Machine Learning Infrastructure Engineer jobs?

Cities near San Jose, CA with the most Machine Learning Infrastructure Engineer job openings:

Infographic showing various Machine Learning Infrastructure Engineer job openings in San Jose, CA as of August 2026, with employment types broken down into 1% As Needed, 76% Full Time, 21% Part Time, and 2% Contract. Highlights an 88% Physical, 2% Hybrid, and 10% Remote job distribution, with an average salary of $148,920 per year, or $71.6 per hour.

$125K - $164K/yr

Full-time

Re-posted 21 days ago


Job description

Job Summary:
MBZUAI is a dedicated research lab focused on advancing AI research and building high-performance computing systems. The Machine Learning Infrastructure Engineer will work on extending and scaling training systems, collaborating with researchers and engineers to tackle challenges in AI development.
Responsibilities:
• Extend distributed training frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
• Implement distributed optimizers from mathematical specs
• Build robust config + launch systems across multi-node, multi-GPU clusters
• Own experiment tracking, metrics logging, and job monitoring for external visibility
• Improve training system reliability, maintainability, and performance
• Extend or modify training frameworks (e.g., DeepSpeed, FSDP) to support new use cases and architectures
• Translate mathematical optimizer specs into distributed implementations
• Create and debug multi-node launch scripts with flexible batch sizes, parallelism strategies, and hardware targets
• Build systems for experiment tracking, job monitoring, and logging usable by collaborators and researchers
• Write production-quality code and tests for ML infra in PyTorch or JAX; ensure reliability and maintainability at scale.
Qualifications:
Required:
• 5+ years of experience in ML systems, infra, or distributed training
• Experience modifying distributed ML frameworks (e.g., DeepSpeed, FSDP, FairScale, Horovod)
• Strong software engineering fundamentals (Python, systems design, testing)
• Proven multi-node experience (e.g., Slurm, Kubernetes, Ray) and debugging skills (e.g., NCCL/GLOO)
• Ability to implement algorithms across GPUs/nodes based on mathematical specs
• Experience working on an ML platform/ infrastructure, and/or distributed inference optimization team
• Experience with large-scale machine learning workloads (strong ML fundamentals)
Preferred:
• Exposure to mixed-precision training (e.g., bf16, fp8) with accuracy validation
• Familiarity with performance profiling, kernel fusion, or memory optimization
• Open-source contributions or published research (MLSys, ICML, NeurIPS)
• CUDA or Triton kernel experience
• Experience with large-scale pre-training
• Experience building custom training pipelines at scale and modifying them for custom needs
• Deep familiarity with training infrastructure and performance tuning
Company:
Official account of Mohamed bin Zayed University of Artificial Intelligence. Dedicated to research, innovation, and empowering brilliant minds in AI. Founded in 2019, the company is headquartered in Abu Dhabi, ARE, with a team of 51-200 employees. The company is currently Growth Stage.