1

Ml Platform Engineer Jobs (NOW HIRING)

Senior ML Platform Engineer

Burbank, CA · On-site

$130.20 - $195.30/hr

Burbank, CA, US, 91505 New York, NY, US, 10036 Research Burbank Full-Time On-Site We are seeking a Sr. ML Platform Engineer who is excited to deploy MLOps products that shape business strategy ...

Toyota Financial Services Enterprise Platforms team is looking for a Senior ML Platform Engineer to design, build, and operationalize an enterprise ML platform on AWS SageMaker Unified Studio. You ...

New

MLOps Platform Engineer (SageMaker)

Plano, TX · On-site

$123.98 - $130.87/hr

Client Enterprise Platforms team is looking for a Senior ML Platform Engineer to design, build, and operationalize an enterprise ML platform on AWS SageMaker Unified Studio. You will migrate the ...

Lead AI/ML Platform Engineer

Plano, TX

$98K - $129K/yr

The Lead AI/ML Platform Engineer will support the Enterprise Platforms team's objective to deliver reliable, secure, and high-performing AI platform capabilities that drive business value at scale.

This role will play a key part in enabling scalable, reliable, and secure ML model development and deployment across our cloud and container platforms. This is a hands-on engineering role requiring ...

Showing results 41-60

Ml Platform Engineer information

See salary details

$33

$63

$94

How much do ml platform engineer jobs pay per hour?

As of Sep 5, 2026, the average hourly pay for ml platform engineer in the United States is $63.95, according to ZipRecruiter salary data. Most workers in this role earn between $50.48 and $73.80 per hour, depending on experience, location, and employer.

What is an ML Platform Engineer?

ML Platform Engineers are specialized software engineers who design, build, and maintain the infrastructure and tools needed to support the development, deployment, and scaling of machine learning models. They bridge the gap between data science and production engineering by automating model training, monitoring, versioning, and serving. Their work enables data scientists to focus on modeling while ensuring that ML solutions are reliable, reproducible, and scalable in real-world environments.

What skills and qualifications are needed to thrive as an ML Platform Engineer?

To thrive as an ML Platform Engineer, you need a strong background in computer science, software engineering, and machine learning concepts, often supported by a degree in a related field. Expertise with cloud platforms (such as AWS, GCP, or Azure), containerization (Docker, Kubernetes), CI/CD pipelines, and knowledge of ML frameworks (TensorFlow, PyTorch) are commonly required. Collaboration, problem-solving, and strong communication skills help you work efficiently with data scientists, engineers, and stakeholders. These skills ensure the development, scalability, and reliability of robust ML infrastructure that empowers teams to deploy and manage models effectively.

How does an ML Platform Engineer typically collaborate with data scientists and software engineers within a company?

ML Platform Engineers work closely with both data scientists and software engineers to streamline the process of developing, deploying, and maintaining machine learning models. They provide the infrastructure and tools necessary for data scientists to build and experiment with models efficiently, while ensuring seamless integration with production systems managed by software engineers. Regular communication, participation in cross-functional meetings, and shared project management tools are common ways teams collaborate. This close collaboration helps to bridge the gap between research and production, ensuring robust, scalable, and reliable ML solutions.

What is the difference between Ml Platform Engineer vs Data Scientist?

AspectML Platform EngineerData Scientist
Required credentialsBachelor's/Master's in CS, Engineering, or related; experience with cloud platformsBachelor's/Master's in Statistics, Math, or CS; strong programming skills
Work environmentBuilds and maintains ML infrastructure, collaborates with engineering teamsAnalyzes data, develops models, and interprets results
Industry usageTech companies, AI startups, enterprises deploying ML systemsResearch institutions, tech firms, data-driven organizations

ML Platform Engineers focus on developing and maintaining the infrastructure that supports machine learning models, while Data Scientists primarily analyze data and build models. Both roles often collaborate but serve different functions within the AI and data ecosystem.

More about Ml Platform Engineer jobs

What cities are hiring for Ml Platform Engineer jobs?

Cities with the most Ml Platform Engineer job openings:

What states have the most Ml Platform Engineer jobs?

States with the most job openings for Ml Platform Engineer jobs include:

What job categories do people searching Ml Platform Engineer jobs look for?

The top searched job categories for Ml Platform Engineer jobs are:

Infographic showing various Ml Platform Engineer job openings in the United States as of August 2026, with employment types broken down into 50% Full Time, 47% Part Time, and 3% Contract. Highlights an 78% Physical, 2% Hybrid, and 20% Remote job distribution, with an average salary of $133,026 per year, or $64 per hour.

Sr. AI/ML Platform Engineer

Advanced Micro Devices

Santa Clara, CA • On-site

$190 - $260/hr

Other

Re-posted 8 days ago


Advanced Micro Devices rating

8.6

Company rating: 8.6 out of 10

Based on 13 frontline employees who took The Breakroom Quiz

26th of 161 rated electronics manufacturers


Job description

WHAT YOU DO AT AMD CHANGES EVERYTHING

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges- striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.

Together, we advance your career.

THE ROLE

We are hiring AI / ML Platform Engineers to build the platform layer that makes AI-for-engineering workflows scalable, reliable, and reproducible. This role focuses on the infrastructure and platform systems that support large-scale agent execution, distributed training and inference, experiment tracking, benchmark automation, artifact management, and GPU cluster utilization.

You will work closely with ML Systems Research Engineers, AI Research Scientists, Applied AI Engineers, and hardware domain experts to operationalize the Blueprint framework across kernel optimization, RTL/PPA optimization, ECO fixing, verification, simulation, and debugging workflows.

This is a platform engineering role, not a pure research role. The focus is to build robust shared systems that allow researchers and engineers to run more experiments, compare results reliably, reduce manual orchestration, and move successful workflows into production engineering use.

THE PERSON

You are a strong systems engineer who enjoys building reliable platforms for AI researchers and applied engineers. You understand distributed systems, ML workloads, GPU infrastructure, experiment management, and production reliability. You can turn messy research workflows into reusable services, APIs, dashboards, job systems, and automation.

You care about reproducibility, observability, performance, and developer experience. You are comfortable working across ML, infrastructure, and hardware/software tooling, and you can partner with research teams without requiring every requirement to be fully specified upfront.

KEY RESPONSIBILITIES
  • Build and operate the shared AI platform for agentic engineering workflows, including job submission, scheduling, orchestration, retries, logging, artifact storage, and experiment tracking.
  • Develop reliable infrastructure for distributed training, distributed inference, batch evaluation, and large-scale agent rollout across GPU clusters.
  • Build platform services for benchmark execution, correctness checking, profiling, regression tracking, and reproducible evaluation.
  • Maintain artifact systems for generated kernels, RTL edits, traces, logs, profiler outputs, benchmark results, simulator outputs, and formal verification artifacts.
  • Support scalable integrations with compilers, ROCm/HIP tooling, profilers, simulators, EDA tools, vLLM, SGLang, and internal engineering systems.
  • Improve GPU cluster utilization, scheduling efficiency, reliability, quota management, and workload isolation.
  • Build dashboards and observability systems for experiment status, resource usage, failure modes, benchmark trends, regression detection, and team productivity.
  • Partner with ML Systems Research Engineers to productionize research workflows for RL systems, inference systems, quantification systems, and evaluation pipelines.
  • Partner with Applied AI Engineers to make Blueprint harnesses reusable across kernel optimization, RTL optimization, verification, firmware, and CPU/GPU performance workflows.
  • Establish platform standards for reproducibility, data retention, run metadata, artifact lineage, access control, and operational reliability.
TECHNICAL FOCUS AREAS
  • Distributed ML platform infrastructure for training, inference, evaluation, and agent execution.
  • GPU cluster scheduling, utilization, reliability, quota management, and multi-user workload isolation.
  • Experiment tracking, artifact management, run lineage, dashboards, and reproducible workflow management.
  • Benchmark and evaluation automation for correctness, performance, regression detection, and reproducibility.
  • Integration with ROCm/HIP, Triton, compilers, profilers, vLLM, SGLang, simulators, formal tools, and EDA flows.
  • Caching and parallelization for expensive feedback loops, including simulator, compiler, verifier, and benchmark workloads.
  • Production-quality APIs, services, workflow engines, and developer tooling for research and engineering teams.
PREFERRED QUALIFICATIONS
  • Strong programming skills in Python and one or more systems languages such as C++, Go, or Rust.
  • Experience building ML platforms, AI infrastructure, distributed systems, workflow orchestration, experiment platforms, GPU cluster infrastructure, or developer platforms.
  • Strong understanding of job scheduling, distributed workloads, logging, monitoring, reliability, storage systems, and production operations.
  • Experience with Kubernetes, Ray, Slurm, workflow engines, containerization, CI/CD, data pipelines, or large-scale compute orchestration.
  • Ability to build reliable services, APIs, dashboards, and developer tools used by researchers and engineers.
  • Strong debugging skills across distributed systems, GPU workloads, storage systems, networking, containers, and production infrastructure.
  • Good collaboration skills with AI researchers, ML systems researchers, applied engineers and hardware domain experts.
PREFERRED EXPERIENCE
  • Experience with GPU platforms, ROCm/HIP, CUDA, profiling, kernel benchmarking, model serving, or distributed training/inference.
  • Experience with vLLM, SGLang, Triton, PyTorch, JAX, Ray, Kubernetes, Slurm, MLflow, Weights & Biases, or similar systems.
  • Experience building experiment tracking systems, artifact stores, workflow engines, benchmark automation platforms, or developer productivity tools.
  • Experience supporting LLM agents, tool-use systems, large-scale sampling, or automated program optimization workflows.
  • Familiarity with compiler, profiler, simulator, formal verification, EDA, firmware, or hardware performance workflows is a strong plus.
  • Experience operating shared GPU clusters or high-performance ML infrastructure in a multi-user research environment.
EDUCATION

Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, Machine Learning, or related field, or equivalent practical experience. Master's preferred; PhD is a plus, especially with work in ML systems, reinforcement learning, distributed systems, GPU computing, or AI infrastructure.

LOCATION: Santa Clara, CA

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD's 'Responsible AI Policy' is available here.

This posting is for an existing vacancy.

#J-18808-Ljbffr

What Advanced Micro Devices employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom