1

Ai Reliability Engineer Jobs in Washington, DC (NOW HIRING)

Site Reliability Engineer

Washington, DC · On-site

$114K - $190K/yr

Manage incident response, root cause analysis, and post-mortem processes for the AI platform ... , DevOps, or production operations. * Extensive experience with cloud-native infrastructure ...

Site Reliability Engineer - Hybrid

Reston, VA · On-site

$59.25 - $78.75/hr

AI/ML: We have certain machine learning projects which the SRE interacts with. So, AI/ML experience is a plus to have. * Previous Fannie Mae experience is a plus. Overall years of experience: 8+ ...

Site Reliability Engineer II

Mclean, VA · On-site

$57.50 - $76.50/hr

As an SRE II, you will help operate and improve the reliability, scalability, and performance of ... Leverage AI-assisted engineering tools and automation platforms to accelerate troubleshooting ...

Site Reliability Engineer II

Mclean, VA · On-site

$103.50 - $150/hr

As an SRE II, you will help operate and improve the reliability, scalability, and performance of ... Leverage AI-assisted engineering tools and automation platforms to accelerate troubleshooting ...

Site Reliability Engineer II

Mclean, VA · On-site

$57.50 - $76.50/hr

As an SRE II, you will help operate and improve the reliability, scalability, and performance of ... Leverage AI-assisted engineering tools and automation platforms to accelerate troubleshooting ...

next page

Showing results 1-20

Ai Reliability Engineer information

See Washington, DC salary details

$69.1K

$133.6K

$159.7K

How much do ai reliability engineer jobs pay per year?

As of Aug 10, 2026, the average yearly pay for ai reliability engineer in Washington, DC is $133,616.00, according to ZipRecruiter salary data. Most workers in this role earn between $116,100.00 and $146,100.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What are popular job titles related to Ai Reliability Engineer jobs in Washington, DC? For Ai Reliability Engineer jobs in Washington, DC, the most frequently searched job titles are:
What job categories do people searching Ai Reliability Engineer jobs in Washington, DC look for? The top searched job categories for Ai Reliability Engineer jobs in Washington, DC are:

Site Reliability Engineer

ManTech

Washington, DC • On-site

$114K - $190K/yr

Other

Medical, Life, Retirement, PTO

Re-posted 20 days ago


ManTech rating

9.0

Company rating: 9.0 out of 10

Based on 14 frontline employees who took The Breakroom Quiz

33rd of 242 rated software companies


Job description

Description & Requirements
Unlock the secrets of intelligence with MANTECH! Join a dynamic team at the forefront of national security, providing advanced solutions to government intelligence agencies. Since 1968, we've been solving the toughest challenges with groundbreaking tech. Explore thrilling projects in Digital Transformation, Cybersecurity, IT, Data Analytics and Software Development. Elevate your career and make a difference. Your adventure begins now-unleash your potential with MANTECH!
***This is for a future opportunity***

MANTECH seeks motivated, career, and customer-oriented Site Reliability Engineer (SRE) for a new initiative. This effort supports the rapid design, deployment, operation, and sustainment of enterprise-scale AI, data, and mission platform capabilities across cloud, edge, and classified operational environment

This role supports the operational reliability, scalability, monitoring, and incident response for the enterprise AI systems. You will focus on operational outcomes and optimizing system performance.

Responsibilities include but are not limited to:

  • Apply core reliability engineering principles to ensure high availability and stability of production systems.

  • Manage incident response, root cause analysis, and post-mortem processes for the AI platform.

  • Implement and optimize observability operations using OpenTelemetry, Prometheus, Grafana, Loki, or Tempo.

  • Oversee capacity planning, performance optimization, and FinOps practices.

  • Define and continuously monitor Service Level Objectives (SLOs) and Service Level Agreements (SLAs).

Minimum Qualifications:

  • Bachelor's degree in Computer Science, Engineering, or a related technical discipline.

  • 5 or more years of experience in Site Reliability Engineering (SRE), DevOps, or production operations.

  • Extensive experience with cloud-native infrastructure, particularly Kubernetes.

  • Deep knowledge of monitoring, alerting, and logging systems.

  • Proven ability to automate operational tasks and reduce toil.

Preferred Qualifications:

  • Hands-on experience with the full observability stack: OpenTelemetry, Prometheus, Grafana, Loki, and Tempo.

  • Experience with FinOps and optimizing cloud resource consumption.

  • Experience supporting high-scale distributed systems in a secure environment.

Clearance Requirements:

  • For onsite work, a TS/SCI clearance with Poly will be required.

Physical Requirements:

  • The person in this position must be able to remain in a stationary position 50% of the time.

  • Frequently communicates with co-workers, management, and customers, which may involve delivering presentations.

  • Constantly operates a computer and other office productivity machinery.


The projected compensation range for this position is $114,600.00-$190,200.00. There are differentiating factors that can impact a final salary/hourly rate, including, but not limited to, Contract Wage Determination, relevant work experience, skills and competencies that align to the specified role, geographic location (For Remote Opportunities), education and certifications as well as Federal Government Contract Labor categories.  In addition, MANTECH invests in its employees beyond just compensation.  MANTECH's benefits offerings include, dependent upon position, Health Insurance, Life Insurance, Paid Time Off, Holiday Pay, short-term and long-term Disability, Retirement and Savings, Learning and Development opportunities, wellness programs as well as other optional benefit elections.

MANTECH considers all qualified applicants for employment without regard to disability or veteran status or any other status protected under any federal, state, or local law or regulation.
If you need a reasonable accommodation to apply for a position with MANTECH, please email us at careers@mantech.com and provide your name and contact information.

What ManTech employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom