1

Ai Reliability Engineer Jobs in Reston, VA (NOW HIRING)

Site Reliability Engineer - CTJ - Poly

Reston, VA · On-site

$59.25 - $78.75/hr

You will gain end-to-end ownership experience across the service lifecycle while helping shape the future of reliability engineering through automation, data-driven operations, and AI-enabled ...

Site Reliability Engineer (US Federal)

Reston, VA · On-site

$59.25 - $78.75/hr

As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we're ... As an SRE, you will help in building a clean, scalable, reliable, and automated services framework.

As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we're ... As an SRE, you will help in building a clean, scalable, reliable, and automated services framework.

As a Fortune 500 company and a leading AI platform for managing people, money, and agents, we're ... As an SRE, you will help in building a clean, scalable, reliable, and automated services framework.

Showing results 41-60

Ai Reliability Engineer information

See Reston, VA salary details

$63.5K

$122.7K

$146.7K

How much do ai reliability engineer jobs pay per year?

As of Aug 17, 2026, the average yearly pay for ai reliability engineer in Reston, VA is $122,734.00, according to ZipRecruiter salary data. Most workers in this role earn between $106,600.00 and $134,200.00 per year, depending on experience, location, and employer.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Reston, VA?

For Ai Reliability Engineer jobs in Reston, VA, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Reston, VA look for?

The top searched job categories for Ai Reliability Engineer jobs in Reston, VA are:

What cities near Reston, VA are hiring for Ai Reliability Engineer jobs?

Cities near Reston, VA with the most Ai Reliability Engineer job openings:

Site Reliability Engineer - CTJ - Poly

Microsoft

Reston, VA • On-site

$59.25 - $78.75/hr

Full-time

Re-posted 13 days ago


Microsoft rating

8.5

Company rating: 8.5 out of 10

Based on 132 frontline employees who took The Breakroom Quiz

79th of 244 rated software companies


Job description

Overview
Help build and operate the trusted platforms that power Microsoft 365's most critical compliance, security, and governance services. As a member of our team, you will work at the intersection of large-scale cloud engineering, service reliability, and operational excellence, helping ensure that enterprise and government customers can rely on Microsoft services every day. You'll collaborate with engineers across Microsoft to solve complex technical challenges, improve resiliency, and drive innovation through automation, modern cloud technologies, and data-driven operations.
As a Site Reliability Engineer, you will help design, operate, and continuously improve large-scale Microsoft 365 and Purview services that support millions of users worldwide. You will partner with software engineers, service owners, and reliability teams to monitor service health, automate operational processes, investigate production issues, and implement engineering solutions that improve availability, performance, security, and customer experience.
This opportunity will allow you to accelerate your cloud engineering expertise, develop deep knowledge of distributed systems and large-scale service operations, and build advanced skills in automation, observability, incident management, and AI-powered operational tooling. You will gain end-to-end ownership experience across the service lifecycle while helping shape the future of reliability engineering through automation, data-driven operations, and AI-enabled solutions.
Microsoft's mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Responsibilities
• You participate in onboarding, code/design reviews, and regular meetings with the engineering teams that develop and manage those products.
• You independently develop code or scripts that automate the performance of repetitive and easily scalable operations processes.
• You design, develop, and maintain telemetry pipelines and monitoring tools that detail operations metrics.
• You develop, test, troubleshoot, and implement changes to optimize code and improve products.
• You respond to incidents during regular on-call rotations.
Qualifications
Required/Minimum Qualifications
  • Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Other Requirements:
Security Clearance Requirements: Candidates must be able to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:
  • The successful candidate must have an active U.S. Government Top Secret Clearance with access to Sensitive Compartmented Information (SCI) based on a Single Scope Background Investigation (SSBI). Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. Failure to maintain or obtain the appropriate U.S. Government clearance and/or customer screening requirements may result in employment action up to and including termination.
  • Clearance Verification: This position requires successful verification of the stated security clearance to meet federal government customer requirements. You will be asked to provide clearance verification information prior to an offer of employment.
  • Microsoft Cloud Background Check: This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.
  • Citizenship & Citizenship Verification: This position requires verification of U.S. citizenship due to citizenship-based legal restrictions. Specifically, this position supports United States federal, state, and/or local United States government agency customer and is subject to certain citizenship-based restrictions where required or permitted by applicable law. To meet this legal requirement, citizenship will be verified via a valid passport, or other approved documents, or verified US government Clearance

Preferred/Additional Qualifications
  • Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Site Reliability Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

What Microsoft employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Microsoft logo

About Microsoft

Sourced by ZipRecruiter

Our infrastructure is comprised of a large global portfolio of more than 100 datacenters and 1 million servers. Our foundation is built upon and managed by a team of subject matter experts working to support services for more than 1 billion customers and 20 million businesses in over 90 countries worldwide. With environmental sustainability and optimization at the forefront of our datacenter design and operations, we continue to grow and evolve as we meet the ever-changing business demands that hold Microsoft as a world-class cloud provider.

Industry

Computer and computer peripheral equipment and software wholesalers

Company size

10,000+ Employees

Headquarters location

Redmond, WA, US

Year founded

1975

Social media