1

Ai Reliability Engineer Jobs in Silver Spring, MD

Fleet Reliability Engineer

Arlington, VA ยท On-site

$117K - $148K/yr

... AI-powered insights that help coast guards, navies, insurers, energy operators, and researchers ... The Role The Fleet Reliability Engineer owns the health, uptime, and long-term reliability of our ...

Staff Cyber Site Reliability Engineer (SRE)

Bethesda, MD ยท On-site

$61 - $81/hr

  • Retirement

Experience with AI/ML and working knowledge of LLMs is a meaningful differentiator. Position Responsibilities As a Staff Cyber SRE, you will: * Own Reliability Engineering: Define and drive ...

Fleet Reliability Engineer

Arlington, VA ยท On-site

$120 - $170/hr

... AI-powered insights that help coast guards, navies, insurers, energy operators, and researchers ... The Role The Fleet Reliability Engineer owns the health, uptime, and long-term reliability of our ...

New

Site Reliability Engineer, Lead

Chantilly, VA ยท On-site

$99K - $225K/yr

  • Medical

  • Life

  • Retirement

  • PTO

As a Lead Site Reliability Engineer (SRE) on our team,you'llbe responsible forensuring the ... Candidate AI Usage Policy AI is a part of our daily work at Booz Allen, and we are committed to the ...

Senior Site Reliability Engineer

Mclean, VA ยท On-site

$57.50 - $76.50/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

As a Senior Site Reliability Engineer, you will play a key role in designing, operating, and ... As AI capabilities continue to evolve, Senior SREs are expected to evaluate, adopt, and promote AI ...

Site Reliability Engineer - Hybrid

Reston, VA ยท On-site

$59.25 - $78.75/hr

AI/ML: We have certain machine learning projects which the SRE interacts with. So, AI/ML experience is a plus to have. * Previous Fannie Mae experience is a plus. Overall years of experience: 8+ ...

Site Reliability Engineer

Chantilly, VA ยท On-site

$62K - $141K/yr

  • Medical

  • Life

  • Retirement

  • PTO

Site Reliability Engineer The Opportunity: Engineering to make a system more resilient and ... Candidate AI Usage Policy AI is a part of our daily work at Booz Allen, and we are committed to the ...

Site Reliability Engineer

Chantilly, VA ยท On-site

$62K - $141K/yr

  • Medical

  • Life

  • Retirement

  • PTO

As a Site Reliability Engineer (SRE) on our team, you'll help the Intelligence Community develop ... Candidate AI Usage Policy AI is a part of our daily work at Booz Allen, and we are committed to the ...

Site Reliability Engineer

Chantilly, VA ยท On-site

$62K - $141K/yr

  • Medical

  • Life

  • Retirement

  • PTO

Site Reliability Engineer The Opportunity: Engineering to make a system more resilient and ... Candidate AI Usage Policy AI is a part of our daily work at Booz Allen, and we are committed to the ...

Site Reliability Engineer II

Mclean, VA ยท On-site

$103.50 - $150/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

As an SRE II, you will help operate and improve the reliability, scalability, and performance of ... Leverage AI-assisted engineering tools and automation platforms to accelerate troubleshooting ...

Site Reliability Engineer II

Mclean, VA ยท On-site

$57.50 - $76.50/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

As an SRE II, you will help operate and improve the reliability, scalability, and performance of ... Leverage AI-assisted engineering tools and automation platforms to accelerate troubleshooting ...

Staff Site Reliability Engineer, Federal (TS/SCI)

Washington, DC ยท On-site

$64.50 - $85.75/hr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Okta secures AI by building the trusted, neutral infrastructure that enables organizations to ... The Federal SRE Team We are looking for an experienced Staff Site Reliability Engineer to join Okta ...

Performance & Reliability Engineer

Washington, DC ยท On-site

$70.50 - $136.70/hr

Here's What You Need: * 3+ years of experience in Site Reliability Engineering * 1+ years of experience with Generative AI * 2 + years of experience with SRE Observability * 3+ years of experience ...

next page

Showing results 1-20

Ai Reliability Engineer information

See Silver Spring, MD salary details

$63.1K

$122K

$145.8K

How much do ai reliability engineer jobs pay per year?

As of Aug 16, 2026, the average yearly pay for ai reliability engineer in Silver Spring, MD is $121,958.00, according to ZipRecruiter salary data. Most workers in this role earn between $106,000.00 and $133,400.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are popular job titles related to Ai Reliability Engineer jobs in Silver Spring, MD?

For Ai Reliability Engineer jobs in Silver Spring, MD, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Silver Spring, MD look for?

The top searched job categories for Ai Reliability Engineer jobs in Silver Spring, MD are:

What cities near Silver Spring, MD are hiring for Ai Reliability Engineer jobs?

Cities near Silver Spring, MD with the most Ai Reliability Engineer job openings:

Fleet Reliability Engineer

Quartermaster AI Inc

Arlington, VA โ€ข On-site

$117K - $148K/yr

Full-time

Posted 4 days ago


Job description

About Quartermaster
Quartermaster is building a live, real-time picture of the world's oceans. Our distributed sensor network turns ordinary civil and commercial vessels into intelligent nodes, delivering HD video, signal intelligence, and AI-powered insights that help coast guards, navies, insurers, energy operators, and researchers monitor, detect, and respond across global waters.
At the center of the network is SmartMastโ„ข - a ruggedized, vessel-mounted sensor package that installs on any vessel over 20 GT. Each unit combines a 360ยฐ/31x optical and infrared camera array (IP68-rated), dual software-defined radios, an onboard AI compute module, a configurable powerpack, and resilient global SATCOM that pushes data to users anywhere in under five seconds. We have deployed 600+ sensors across more than 25 countries, and our fleet has traveled over 10 million nautical miles.
As the fleet scales, keeping every deployed SmartMast healthy, connected, and producing reliable data at sea is mission-critical. We are hiring our first dedicated Fleet Reliability Engineer to own that outcome.
The Role
The Fleet Reliability Engineer owns the health, uptime, and long-term reliability of our deployed SmartMast hardware fleet. You will be the person who knows - at any moment - how many units are online, which are degrading, why units fail, and what we are doing about it. You will turn fleet telemetry into action: catching failures before they take a unit offline, driving root-cause analysis on the ones that slip through, and closing the loop with hardware design, firmware, and field-service teams so the same failure never recurs.
This is a hands-on, high-ownership role suited to a senior engineer who is comfortable operating at the intersection of hardware, embedded systems, data, and harsh-environment field operations. You will build the reliability program from the ground up: the metrics, the monitoring, the failure-tracking process, and the maintenance and RMA workflows that let the fleet scale from hundreds to thousands of units.
What You'll Do
Fleet health & monitoring
  • Own fleet-wide reliability metrics - uptime, availability, MTBF, MTTR, data-yield, and failure rates by component and by deployment environment - and report them to engineering and leadership.
  • Build and refine dashboards, alerting, and telemetry pipelines that surface degrading units (power, thermal, connectivity, camera, radio, compute) before they go offline.
  • Define what "healthy" means for each subsystem and set the thresholds that trigger proactive intervention.
Failure analysis & continuous improvement
  • Lead root-cause analysis (RCA) on field failures, from telemetry forensics through physical teardown of returned units.
  • Maintain the fleet failure database and drive FMEA, reliability growth tracking, and corrective/preventive action (CAPA) to closure.
  • Close the loop with hardware, firmware, and manufacturing teams - translating field failures into design-for-reliability, component-selection, and firmware changes.
Field service, maintenance & logistics
  • Define preventive maintenance schedules, spares strategy, and the RMA / repair-and-return process for a globally distributed fleet.
  • Update installation, diagnostic, and field-repair procedures and troubleshooting guides used by internal technicians and partner crews.
  • Support field deployments and complex repairs directly, including periodic travel to vessels, ports, and installation sites.
Reliability engineering & scale
  • Work with the hardware team to establish environmental and life-test protocols (vibration, salt-fog/corrosion, thermal, ingress, power) to qualify hardware and predict field life before deployment.
  • Feed reliability requirements and acceptance criteria into new hardware revisions and supplier qualification.
  • Design the reliability processes and tooling so they scale as the fleet grows into the thousands of units.
Minimum Requirements
  • Bachelor's degree in Electrical, Mechanical, Systems, Reliability, or a related engineering discipline - or equivalent hands-on experience.
  • 5+ years of engineering experience with deployed electro-mechanical hardware, at least 2 of which are in reliability, sustaining/field engineering, or hardware operations for a fielded product.
  • Demonstrated ownership of hardware reliability outcomes for a fleet or installed base - you have been directly responsible for uptime, failure rates, or MTBF/MTTR of real hardware in the field.
  • Hands-on proficiency with root-cause analysis methods (8D, 5-Whys, fishbone) and reliability tools such as FMEA, fault-tree analysis, and CAPA.
  • Practical experience diagnosing electro-mechanical systems using telemetry/logs, bench instruments, and physical teardown.
  • Data fluency: able to query, analyze, and visualize fleet telemetry using SQL and Python (or equivalent) to find trends and drive decisions.
  • Working knowledge of electronics, power systems, and mechanical enclosures, and the failure modes of hardware operating in harsh outdoor environments.
  • Willingness and ability to travel periodically to field sites (vessels, ports, installation locations), including occasional international travel.
  • Must be legally authorized to work in the United States and able to satisfy any customer- or contract-driven eligibility requirements associated with government and maritime-security work.
Preferred Qualifications
  • Experience with hardware deployed in marine, maritime, offshore, automotive, aerospace/defense, satellite, telecom, or other remote/harsh-environment fleets.
  • Familiarity with IP-rated enclosures, corrosion and salt-fog effects, marine power systems, and environmental qualification (e.g., IEC 60529, MIL-STD-810, IEC 60068).
  • Experience with connected/IoT or edge devices: remote diagnostics, OTA firmware updates, and interpreting embedded-system and connectivity (SATCOM/cellular) telemetry.
  • Exposure to camera/optical systems, RF/software-defined radios, batteries, or edge-AI compute hardware.
  • Background building a reliability or sustaining-engineering function from scratch at a hardware startup or scaling operation.
  • ASQ Certified Reliability Engineer (CRE) or comparable credential.
What Success Looks Like
  • Fleet uptime and data-yield are measured, trending up, and visible to the whole company.
  • Failures are caught proactively from telemetry rather than reported by customers.
  • Every significant field failure has a documented root cause and a closed corrective action.
  • Reliability feedback is shaping each new SmartMast hardware revision, and the reliability program scales cleanly as the fleet grows.
Why Join Us
You will build the reliability backbone of a fast-growing maritime intelligence network and see your work sail across the world's oceans. Our leadership brings deep experience across the Navy, DARPA, Anduril, and Scale AI, and we are backed to scale a fleet that is already operating in 25+ countries. If you want to own hardware reliability end-to-end and watch your impact compound with every vessel that joins the network, we would like to talk.