1

Mlops Manager Jobs (NOW HIRING)

This is a hands-on technical leadership role, not a management position; you will be a primary ... Own the technical direction for the MLOps platform - define subsystem interfaces, drive ...

Architect and manage the cloud infrastructure supporting the MLOps platform, leveraging infrastructure-as-code (IaC) tools like Terraform. Optimize for scalability, security, cost-effectiveness, and ...

Architect and manage the cloud infrastructure supporting the MLOps platform, leveraging infrastructure-as-code (IaC) tools like Terraform. Optimize for scalability, security, cost-effectiveness, and ...

MLOps Platform Engineer Location: Reston VA - In person interviews so need Local In EAST coast only ... Container & Kubernetes Workloads · Design and manage EKS workloads supporting containerized ML ...

MLOps Platform Engineer Location: Reston VA Required Qualifications • 3+ years of hands-on ... Preferred Qualifications • Experience managing Data Analytics Platforms / Tools (e.g., Domino ...

Establish and maintain MLOps practices, including automated training, deployment, monitoring, retraining, and performance management. Ensure AI solutions are reliable, scalable, secure, and optimized ...

This is a hands-on technical leadership role, not a management position; you will be a primary ... Own the technical direction for the MLOps platform - define subsystem interfaces, drive ...

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a MLOps Engineer based in Netherlands. Join a high-impact engineering ...

New

San Francisco, CA, USA (Hybrid/Remote) Job Type: Full-Time About the Role We are seeking an experienced MLOps Engineer to build and manage scalable machine learning infrastructure, automate model ...

MLOps Platform Engineer Location: Reston VA Required Qualifications · 3+ years of hands-on ... managing CI/CD pipelines (GitLab or equivalent). · Familiarity with machine learning workflows ...

We are seeking a senior MLOps Architect to design and scale a modern ML and Generative AI platform ... Managed services (e.g., SageMaker endpoints, Bedrock-style APIs) * Containerized custom inference ...

next page

Showing results 1-20

Mlops Manager information

What engineer makes $500,000 a year?

Senior machine learning engineers and MLOps managers with extensive experience, advanced skills in cloud platforms, and expertise in deploying scalable AI systems can earn $500,000 or more annually. High compensation often reflects leadership roles, specialized knowledge, and working in high-demand industries or companies with competitive benefits.

What is the difference between Mlops Manager vs Data Scientist?

AspectMlops ManagerData Scientist
Required CredentialsBachelor's/Master's in CS, Engineering, or related; certifications in cloud platforms or MLOps toolsBachelor's/Master's in CS, Statistics, or related; certifications in data analysis or machine learning
Work EnvironmentCollaborates with engineering, DevOps, and data teams to deploy and maintain ML systemsAnalyzes data, builds models, and provides insights to inform business decisions
Employer & Industry UsageTech companies, AI startups, enterprises implementing ML pipelinesResearch institutions, tech firms, finance, healthcare, and marketing sectors

The Mlops Manager focuses on deploying, maintaining, and optimizing machine learning systems within an organization, working closely with engineering and DevOps teams. In contrast, a Data Scientist primarily analyzes data, develops models, and provides insights. While both roles require knowledge of machine learning, the Mlops Manager emphasizes operationalizing ML solutions, whereas the Data Scientist emphasizes data analysis and modeling.

What are the key skills and qualifications needed to thrive as an MLOps Manager, and why are they important?

To thrive as an MLOps Manager, you need expertise in machine learning, software engineering, and DevOps practices, often backed by a degree in computer science or a related field. Familiarity with tools like Docker, Kubernetes, CI/CD pipelines, cloud platforms (AWS, Azure, GCP), and certifications such as AWS Certified Machine Learning or Google Cloud Professional ML Engineer are highly beneficial. Strong leadership, problem-solving, and cross-functional communication skills help manage teams and bridge the gap between data science and IT operations. These abilities are crucial for ensuring reliable, scalable, and efficient deployment of machine learning solutions in production environments.

What is a $900000 AI job?

A $900,000 AI job typically refers to a high-level position in artificial intelligence, such as an AI executive, senior machine learning engineer, or AI research director, often requiring advanced skills, extensive experience, and leadership responsibilities. These roles may involve overseeing AI strategy, developing complex models, and managing teams, with compensation reflecting the seniority and impact of the position.

Will MLE be replaced by AI?

Machine Learning Engineers (MLEs) design, develop, and maintain AI systems, and their role is unlikely to be fully replaced by AI. Instead, AI tools can augment their work by automating routine tasks, allowing MLEs to focus on complex problem-solving, model optimization, and system integration. Continuous learning and expertise in AI frameworks and programming are essential for MLEs to stay relevant in evolving technological environments.

What are some common challenges an MLOps Manager faces when integrating machine learning models into production environments?

MLOps Managers often encounter challenges such as ensuring seamless collaboration between data science and engineering teams, managing model versioning, and maintaining reliable deployment pipelines. Balancing rapid experimentation with the need for robust, scalable, and secure production systems can be complex. Additionally, monitoring model performance post-deployment and handling data drift or model degradation are ongoing responsibilities. Effective communication and establishing standardized processes are key to overcoming these challenges and ensuring successful model operations.

Is MLOps in high demand?

MLOps managers are in high demand due to the increasing adoption of machine learning and AI across industries. Organizations seek professionals skilled in deploying, monitoring, and maintaining ML models using tools like Kubernetes, Docker, and cloud platforms, making MLOps a rapidly growing field with strong job prospects.

What are MLOps Managers?

MLOps Managers are professionals responsible for overseeing the deployment, operation, and scaling of machine learning models in production environments. They coordinate teams to ensure seamless collaboration between data scientists, engineers, and IT staff, facilitating the automation of machine learning workflows. Their role involves managing infrastructure, optimizing processes for model monitoring and maintenance, and ensuring compliance with organizational and industry standards. MLOps Managers play a key role in bridging the gap between model development and operationalization, ensuring that machine learning solutions are reliable, reproducible, and scalable.
More about Mlops Manager jobs
What cities are hiring for Mlops Manager jobs? Cities with the most Mlops Manager job openings:
What are the most commonly searched types of Mlops jobs? The most popular types of Mlops jobs are:
What states have the most Mlops Manager jobs? States with the most job openings for Mlops Manager jobs include:
Infographic showing various Mlops Manager job openings in the United States as of July 2026, with employment types broken down into 86% Full Time, 13% Part Time, and 1% Contract. Highlights an 94% Physical, 2% Hybrid, and 4% Remote job distribution.
Staff MLOps Engineer

Other

Posted 25 days ago


Job description

JOB SUMMARY

Apptronik is seeking a Staff MLOps Engineer to own the technical direction of our MLOps platform - the system of record for datasets, experiments, model artifacts, and serving paths that connects teleoperation data collection on one side to deployed autonomy on Apollo on the other. In this role, you will set the architecture for the platform layer above the training cluster: dataset lifecycle, experiment tracking, model registry, evaluation harnesses, and the serving / packaging path that delivers trained policies to robots in the field. You will lead by influence across MLOps, Autonomy, Data Platform, and TeleOp - establishing the standards, contracts, and tooling that turn one-off research code into a repeatable, auditable pipeline from data to deployed model. This is a hands-on technical leadership role, not a management position; you will be a primary contributor while mentoring the engineers around you, and partnering closely with the Training Infrastructure engineer who owns the cluster layer beneath the platform.

ESSENTIAL DUTIES AND RESPONSIBILITIES

Platform Architecture & Ownership

  • Technical Direction: Own the technical direction for the MLOps platform - define subsystem interfaces, drive architecture decisions, and establish engineering standards for how datasets, experiments, and models move through Apptronik's systems.
  • Cross-Team Authority: Serve as the primary technical point of contact for Autonomy, Data Platform, and TeleOp on all matters of model lifecycle and platform contracts.

Dataset Lifecycle & Versioning

  • Versioning & Lineage: Design and operate the dataset layer end-to-end - versioning, lineage, splits, and labeling-integration handoff.
  • Reproducibility: Ensure every trained model can be traced back to the exact data and code that produced it.

Model Registry & Artifact Management

  • Registry: Build and operate a first-class model registry - versioned artifacts, metadata, evaluation results, lineage, and approval workflows.
  • Promotion Path: Define the promotion path from "trained" to "qualified" to "deployed to robot."

Evaluation & Qualification Harnesses

  • Automated Evaluation: Define the offline benchmarks, simulation rollouts, and policy-gating harnesses that any model must pass before reaching Apollo.
  • Metrics Framework: Develop the metrics framework that the autonomy team trusts to gate releases.

Serving, Packaging & Deployment to Robot

  • On-Robot Path: Own the path from registered model to running inference on Apollo - packaging (ONNX, TensorRT, torch.compile), versioning on-robot, rollback, and observability of deployed policy behavior.
  • Telemetry Seam: Coordinate with Connect and Data Platform on the deploy-and-telemetry seam back from the fleet.

Mentorship & Cross-Functional Leadership

  • Mentorship: Mentor mid-level and senior engineers on the MLOps team through code review, design review, and direct collaboration.
  • Influence: Partner with the Training Infrastructure engineer on the cluster/platform contract, and influence research workflows across Autonomy to standardize on the platform's primitives.

SKILLS AND REQUIREMENTS

  • Deep proficiency in Python and at least one systems-level language (Go, Rust, or C++), with demonstrated ability to make and defend architectural tradeoffs in production ML platforms
  • Proven experience owning and delivering an MLOps platform end-to-end - dataset lifecycle, experiment tracking, model registry, evaluation, and serving - at a company that ships models to production
  • Expertise across the model lifecycle: dataset versioning (DVC, LakeFS, Delta, or equivalent), experiment tracking (MLflow, W&B, Determined), model registry, and policy serving
  • Strong background designing service-oriented systems on Kubernetes; comfortable with the contract between platform APIs and underlying compute infrastructure
  • Experience defining evaluation and qualification frameworks for ML models where the cost of a regression is high (robotics, safety-critical, or production-customer-facing)
  • Experience leading technical projects end-to-end: architecture, implementation, validation, and iteration
  • Demonstrated ability to lead by influence across teams - setting standards that other engineers adopt voluntarily, and mentoring engineers around you
  • Proficiency with cloud infrastructure (AWS, GCP, or Azure), Docker, Git, and modern CI/CD workflows

EDUCATION and/or EXPERIENCE

  • Master's degree in Computer Science, Machine Learning, or a related technical field preferred; Bachelor's considered with exceptional experience.
  • 8+ years of professional software engineering experience in ML platforms or related infrastructure, OR 4+ years of direct, hands-on experience owning an MLOps platform that shipped models to production.

 Preferred Qualifications:

  • Experience deploying ML models to edge or embedded targets (on-device inference, ONNX Runtime, TensorRT, robot fleets)
  • Experience with RL training and evaluation infrastructure for embodied agents (rollout workers, replay buffers, sim-eval harnesses)
  • Familiarity with humanoid robotics, dexterous manipulation, or teleoperation data domains
  • Experience with simulation-in-the-loop evaluation (IsaacSim, MuJoCo, or equivalent)
  • Familiarity with policy gating, shadow deployments, or staged rollout strategies for autonomy
  • Open-source contributions to MLOps platform tooling (MLflow, BentoML, KServe, Ray Serve, etc.)

PHYSICAL REQUIREMENTS

  • Prolonged periods of sitting at a desk and working on a computer
  • Must be able to lift 15 pounds at times
  • Vision to read printed materials and a computer screen
  • Hearing and speech to communicate