1

Senior Machine Learning Ops Engineer Jobs in Reston, VA

Senior Machine Learning Engineer

Arlington, VA · On-site

$120K - $165K/yr

We are seeking an experienced Senior Machine Learning Engineer to join our AI/ML team and build the infrastructure that powers the development, evaluation, deployment, and continuous improvement of ...

New

Machine Learning/AI Engineer Location: Hybrid in Vienna, VA or Remote Pay Rate: Open to Both W2 and ... Knowledge of Machine Learning Ops and CI/CD tools for automation of build, test, and deploy models ...

Senior Machine Learning Engineer

Washington, DC · On-site

$118K - $162K/yr

Protagonist is looking for a Senior Machine Learning Engineer who builds production-ready systems for mission-critical work. You bring strong engineering discipline, systems thinking, and a bias ...

Senior Machine Learning Engineer

Reston, VA · Hybrid

$108K - $149K/yr

Position Overview Pantheon Data is seeking a Senior Machine Learning Engineer to design, build, and operate production AI systems for federal clients - including hybrid retrieval-augmented generation ...

Senior Machine Learning Engineer

Reston, VA · On-site

$140K - $200K/yr

Position Overview Pantheon Data is seeking a Senior Machine Learning Engineer to design, build, and operate production AI systems for federal clients - including hybrid retrieval-augmented generation ...

Senior Machine Learning Engineer

Reston, VA · Remote

$140K - $200K/yr

Position OverviewPantheon Data is seeking a Senior Machine Learning Engineer to design, build, and operate production AI systems for federal clients - including hybrid retrieval-augmented generation ...

Showing results 21-40

Senior Machine Learning Ops Engineer information

See Reston, VA salary details

$61.9K

$131.7K

$190.9K

How much do senior machine learning ops engineer jobs pay per year?

As of Sep 10, 2026, the average yearly pay for senior machine learning ops engineer in Reston, VA is $131,664.00, according to ZipRecruiter salary data. Most workers in this role earn between $108,700.00 and $149,300.00 per year, depending on experience, location, and employer.

What is a senior machine learning ops engineer?

Senior Machine Learning Ops (MLOps) Engineers are experienced professionals who design, build, and maintain the infrastructure and tools needed to deploy, monitor, and scale machine learning models in production environments. They work at the intersection of data science, software engineering, and DevOps to ensure ML models are robust, reliable, and secure. Their responsibilities often include automating model training pipelines, managing cloud resources, implementing CI/CD for ML, and ensuring model reproducibility. Senior MLOps Engineers also mentor junior staff and help define best practices for the organization’s ML workflow.

What are the key skills and qualifications needed to thrive as a senior machine learning ops engineer?

To thrive as a Senior Machine Learning Ops Engineer, you need expertise in machine learning, software engineering, cloud platforms, and experience with CI/CD pipelines, often supported by a computer science degree or equivalent experience. Proficiency with tools like Docker, Kubernetes, TensorFlow, PyTorch, and cloud services such as AWS, GCP, or Azure is typically required, along with familiarity with MLOps frameworks. Strong problem-solving, collaboration, and communication skills help you work effectively with cross-functional teams and manage complex ML model deployments. These skills are essential to ensure reliable, scalable, and efficient deployment of machine learning models in production environments.

What are some common challenges faced by senior machine learning ops engineers when deploying models to production?

Senior Machine Learning Ops Engineers often encounter challenges such as ensuring model reproducibility, managing model versioning, and automating deployment pipelines for scalability. Another key challenge is monitoring model performance and data drift in production, which requires robust logging and alerting systems. Collaborating closely with data scientists, software engineers, and IT teams is essential to address these challenges and maintain a stable, efficient ML infrastructure.

What is the difference between Senior Machine Learning Ops Engineer vs Data Engineer?

AspectSenior Machine Learning Ops EngineerData Engineer
CredentialsExperience with ML frameworks, cloud platforms, scripting, and DevOps toolsStrong SQL, ETL, database, and programming skills, often with cloud experience
Work EnvironmentFocus on deploying, monitoring, and maintaining ML models in productionDesigning and building data pipelines and infrastructure for data processing
Industry UsageCommon in AI/ML-focused companies, tech firms, and data-driven organizationsWidespread across industries for data management and analytics

While both roles involve working with data and cloud platforms, the Senior Machine Learning Ops Engineer specializes in deploying and maintaining machine learning models, whereas the Data Engineer focuses on building data pipelines and infrastructure. Understanding these distinctions helps in choosing the right career path or job search focus.

What are popular job titles related to Senior Machine Learning Ops Engineer jobs in Reston, VA?

For Senior Machine Learning Ops Engineer jobs in Reston, VA, the most frequently searched job titles are:

What job categories do people searching Senior Machine Learning Ops Engineer jobs in Reston, VA look for?

The top searched job categories for Senior Machine Learning Ops Engineer jobs in Reston, VA are:

Infographic showing various Senior Machine Learning Ops Engineer job openings in Reston, VA as of September 2026, with employment types broken down into 1% Internship, 1% As Needed, 76% Full Time, 20% Part Time, and 2% Contract. Highlights an 86% Physical, 2% Hybrid, and 12% Remote job distribution, with an average salary of $131,664 per year, or $63.3 per hour.

Senior Machine Learning Engineer

Arlington, VA • On-site

$120K - $165K/yr

Full-time

Posted yesterday

New


Job description

Company Description
Air is the leader in Enterprise Readiness. Our mission is to establish readiness as a real-time condition that is continuously achieved. Today, a dangerous Readiness Gap exists between what the front line needs and what is delivered. Our AI-native platform, Air Enterprise Readiness, aligns development, production, delivery, and sustainment into one coordinated execution system for government agencies and industrial suppliers. By revealing true capacity, exposing real constraints, coordinating resources, and executing at the speed of operational demands, the front line gets what it needs to succeed.
Job Description

We are seeking an experienced Senior Machine Learning Engineer to join our AI/ML team and build the infrastructure that powers the development, evaluation, deployment, and continuous improvement of our language models and AI systems.

As our AI capabilities expand, we need robust infrastructure for moving models from experimentation into production. This role will own critical parts of that lifecycle, including LLMOps, fine-tuning infrastructure, model evaluation, dataset pipelines, experiment management, model serving, and production observability.

In order to do this job well: This is an engineering-heavy ML role. You will build platforms and infrastructure that allow AI engineers and researchers to rapidly experiment with models, datasets, and training techniques while maintaining the reproducibility, scalability, and reliability required for production systems.

You will work across the full model lifecycle - from dataset creation and experimentation through training, evaluation, deployment, monitoring, and iteration.

This role is a full-time position based in our Pittsburgh, PA office or open to Remote Opportunities.

This role may require up to 25% travel, including periodic travel to our Pittsburgh, PA and Arlington, VA offices for team collaboration, planning activities, and in-person meetings.

Scope of Responsibilities
  • Design and build LLMOps infrastructure supporting the development, evaluation, deployment, and continuous improvement of production language models.
  • Build scalable training and fine-tuning infrastructure for commercial and open-weight language models.
  • Develop pipelines supporting supervised fine-tuning, parameter-efficient fine-tuning, preference optimization, and other post-training techniques.
  • Build infrastructure for distributed training and GPU-accelerated ML workloads.
  • Develop data pipelines for training, fine-tuning, evaluation, and synthetic data generation.
  • Build systems for dataset versioning, lineage, quality validation, transformation, and reproducible experimentation.
  • Develop experiment management infrastructure that enables engineers to compare models, datasets, hyperparameters, prompts, and training techniques.
  • Build automated evaluation pipelines that determine whether new models or model versions are ready for production deployment.
  • Design model registries, artifact management, versioning, and promotion workflows across development and production environments.
  • Build and operate scalable model-serving and inference infrastructure for open-weight and fine-tuned models.
  • Develop abstractions that allow product and AI engineering teams to use multiple models and inference providers without tightly coupling applications to a single model or vendor.
  • Build observability for model training and inference, including metrics, tracing, logging, resource utilization, model quality, latency, throughput, and cost.
  • Optimize training and inference workloads for GPU utilization, throughput, latency, reliability, and infrastructure cost.
  • Build automated workflows for model deployment, rollback, canarying, and production validation.
  • Investigate model and infrastructure failures across data pipelines, training jobs, inference services, distributed systems, and production environments.
  • Evaluate emerging models, training techniques, inference frameworks, and ML infrastructure and determine where they can improve our production systems.
  • Partner closely with AI engineers building agentic systems to provide the model, evaluation, and training infrastructure required to continuously improve those systems.
Qualifications
  • U.S. Citizenship is required
Required Skills: 
  • 5+ years of experience building production machine learning systems, ML infrastructure, distributed systems, or similar technical systems.
  • Deep experience designing, building, and operating production ML infrastructure or ML platforms.
  • Experience building infrastructure for training, fine-tuning, evaluating, deploying, and monitoring large language models or other large-scale deep learning models.
  • Experience with LLM fine-tuning and post-training workflows, including techniques such as supervised fine-tuning, LoRA/QLoRA or other parameter-efficient approaches, and preference optimization.
  • Strong understanding of the modern LLM lifecycle, including data preparation, training, evaluation, model artifacts, deployment, inference, monitoring, and iteration.
  • Experience building reproducible ML pipelines involving dataset versioning, experiment tracking, model versioning, and automated evaluation.
  • Experience building and operating production GPU infrastructure across AWS, GCP, Azure, or dedicated GPU providers, including training and/or inference workloads.
  • Strong understanding of distributed systems and the challenges involved in running computationally intensive ML workloads at scale.
  • Strong programming experience in Python and experience building production-quality software.
  • Deep experience with containers, Kubernetes, and cloud platforms such as AWS, GCP, or Azure.
  • Experience designing scalable APIs, services, asynchronous workloads, and data-processing pipelines.
  • Strong understanding of observability and operational reliability for production ML systems.
  • Comfortable debugging failures across training code, datasets, models, GPUs, distributed systems, and cloud infrastructure.
  • Able to move between ML experimentation and infrastructure engineering, understanding the needs of researchers and AI engineers while building systems that make those workflows scalable and reproducible.
  • Comfortable working in a rapidly evolving field where tooling, models, and best practices change quickly.
  •  

Desired Skills: 

  • Current possession of a U.S. security clearance, or the ability to obtain one with our sponsorship
  • Experience in or exposure to the nuances of a startup or other entrepreneurial environment
  • Experience building secure code execution environments or sandboxes for AI agents.
  • Experience with multi-agent architectures, agent-to-agent communication, or distributed agent execution.
  • Experience with fine-tuning, post-training, reinforcement learning, or synthetic data generation.
  • Experience building AI observability, tracing, and debugging infrastructure.
  • Experience optimizing inference latency, throughput, GPU utilization, or model-serving costs.
  • Experience with AI security, adversarial testing, or securing agentic systems.
  • Experience working in government, defense, or other mission-critical environments.
We firmly believe that past performance is the best indicator of future performance.  If you thrive while building solutions to complex problems, are a self-starter, and are passionate about making an impact in global security, we're eager to hear from you.
 
Air is an Equal Opportunity Employer.  All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans status or any other characteristic protected by law.