1

Service Reliability Engineer Jobs in Washington (NOW HIRING)

YOUR FUTURE ROLE We are looking for a Senior Reliability Engineer (SRE) within the Shared Management Services (SMS) group in the Technology and Engineering unit of SAP Sovereign Cloud organization.

Fleet Reliability Engineer

Arlington, VA · On-site

$118K - $148K/yr

Field service, maintenance & logistics * Define preventive maintenance schedules, spares strategy ... Reliability engineering & scale * Work with the hardware team to establish environmental and life ...

Site Reliability Engineer IV

Sterling, VA · On-site

$56.50 - $75/hr

... elevate trust, service, and innovation. Success here requires flexibility in a fast-paced ... Position: Site Reliability Engineer IV Experience: 9-12 Years Location: Bangalore (Ecospace ...

Reliability Engineer

Mclean, VA · On-site

$103K - $130K/yr

The Reliability Engineer will act generally as a member of a design, analysis or review team on ... Positions governed by a Collective Bargaining Agreement (CBA), the McNamara-O'Hara Service Contract ...

Reliability Engineer

Mclean, VA · On-site

$105K - $132K/yr

The Reliability Engineer will act generally as a member of a design, analysis or review team on ... Positions governed by a Collective Bargaining Agreement (CBA), the McNamara-O'Hara Service Contract ...

Reliability Engineer

Mclean, VA · On-site

$103K - $130K/yr

The Reliability Engineer will act generally as a member of a design, analysis or review team on ... Positions governed by a Collective Bargaining Agreement (CBA), the McNamara-O'Hara Service Contract ...

Site Reliability Engineer II - CTJ - Poly

Reston, VA · On-site

$59.25 - $78.75/hr

You will partner with software engineers, service owners, and reliability teams to monitor service health, automate operational processes, investigate production issues, and implement engineering ...

Site Reliability Engineer

Washington, DC · On-site

$112K - $179K/yr

The SRE will drive automation initiatives, observability improvements, and incident response ... Define and manage Service Level Objectives (SLOs) and Service Level Indicators (SLIs). * Support ...

Site Reliability Engineer

Washington, DC · On-site

$112K - $179K/yr

The SRE will drive automation initiatives, observability improvements, and incident response ... Define and manage Service Level Objectives (SLOs) and Service Level Indicators (SLIs). * Support ...

DevOps Engineer

Washington, DC · Hybrid

$59.75 - $81.75/hr

Site Reliability & Service Management Integration * Partner with SRE, network, and platform teams to ensure configuration data supports reliability engineering, availability targets, and operational ...

DevOps Engineer

Washington, DC · On-site

$59.75 - $81.75/hr

Site Reliability & Service Management Integration * Partner with SRE, network, and platform teams to ensure configuration data supports reliability engineering, availability targets, and operational ...

Showing results 41-60

Service Reliability Engineer information

See Washington salary details

$69.1K

$133.6K

$159.7K

How much do service reliability engineer jobs pay per year?

As of Sep 13, 2026, the average yearly pay for service reliability engineer in Washington is $133,616.00, according to ZipRecruiter salary data. Most workers in this role earn between $116,100.00 and $146,100.00 per year, depending on experience, location, and employer.

What is a service reliability engineer?

Service Reliability Engineers (SREs) are IT professionals who apply software engineering principles to infrastructure and operations problems. Their main goal is to ensure that services are reliable, scalable, and highly available by automating processes, monitoring system performance, and responding to incidents. SREs work closely with development and operations teams to design, build, and maintain robust systems, often using code to manage infrastructure. They also focus on improving system reliability through monitoring, incident response, and post-incident analysis.

How does a service reliability engineer typically collaborate with development and operations teams to improve service uptime?

Service Reliability Engineers (SREs) work closely with both development and operations teams to ensure systems are highly available and resilient. They often participate in incident response, conduct post-incident reviews, and help implement automation to reduce manual intervention. Regular collaboration includes reviewing application changes, contributing to infrastructure design, and sharing best practices for monitoring and alerting. This cross-functional teamwork helps to quickly identify potential issues and proactively enhance system reliability.

What are the key skills and qualifications needed to thrive as a service reliability engineer, and why are they important?

To thrive as a Service Reliability Engineer, you need a solid background in systems administration, networking, coding (often in Python or Go), and experience with cloud infrastructure, typically supported by a degree in computer science or a related field. Familiarity with monitoring tools (like Prometheus), CI/CD pipelines, automation frameworks, and certifications such as AWS Certified DevOps Engineer are highly valued. Strong problem-solving abilities, collaboration, and effective communication skills help you proactively address issues and work well within cross-functional teams. These skills ensure system reliability, quick incident recovery, and the seamless delivery of high-availability services.

What is the difference between Service Reliability Engineer vs Site Reliability Engineer?

AspectService Reliability EngineerSite Reliability Engineer
CredentialsTypically requires experience in software engineering, cloud platforms, and monitoring toolsSimilar credentials, often with a focus on software development and systems engineering
Work EnvironmentWorks closely with development and operations teams to ensure service reliabilityWorks on maintaining and improving system reliability, often in cloud or data center environments
Industry UsageCommon in tech companies focusing on service uptime and customer experienceWidely used in tech, especially in cloud and large-scale infrastructure companies

Both roles focus on ensuring system reliability, often requiring similar skills and certifications. The main difference lies in terminology preference and specific organizational focus, but they generally perform comparable functions in maintaining high service availability.

How much do service reliability engineers get paid?

Service Reliability Engineers typically earn a median annual salary between $90,000 and $130,000, depending on experience, location, and industry. Senior roles or those with specialized skills in cloud platforms and automation tools can command higher compensation, often exceeding $150,000 annually.

What are popular job titles related to Service Reliability Engineer jobs in Washington?

For Service Reliability Engineer jobs in Washington, the most frequently searched job titles are:

What job categories do people searching Service Reliability Engineer jobs in Washington look for?

The top searched job categories for Service Reliability Engineer jobs in Washington are:

Infographic showing various Service Reliability Engineer job openings in Washington as of August 2026, with employment types broken down into 1% As Needed, 79% Full Time, 16% Part Time, and 4% Contract. Highlights an 88% Physical, 1% Hybrid, and 11% Remote job distribution, with an average salary of $133,616 per year, or $64.2 per hour.

DevOps Site Reliability Engineer (SRE)

Washington, DC • On-site

IT Veterans
11 - 50 employees

$64.50 - $85.75/hr

Other

Re-posted 20 days ago


Job description

DevOps Site Reliability Engineer (SRE)

Location: Washington, D.C.

Clearance Required: TS/SCI

Position Overview

IT Veterans is seeking a DevOps Site Reliability Engineer (SRE) to support the reliability, performance, and operational stability of a mission-critical enterprise platform. This position plays a vital role in ensuring continuous availability across multi-cloud environments while supporting software deployments, infrastructure monitoring, incident response, and system automation.

The ideal candidate will help bridge development and operations by implementing reliable deployment practices, building robust monitoring capabilities, and rapidly responding to production issues to maintain the required 99.9% platform availability for critical Department of Defense (DoD) systems.

Key Responsibilities:
  • Monitor the health and performance of enterprise infrastructure through continuous system monitoring and automated telemetry to support the required 99.9% platform uptime.
  • Participate in an on-call rotation and respond to major incidents or platform outages within one hour of notification, executing rapid troubleshooting and system stabilization activities.
  • Develop, maintain, and enhance automation scripts and internal tools that streamline diagnostics, health checks, and routine operational tasks.
  • Design and maintain dashboards that provide real-time visibility into platform health, including uptime, API performance, incident status, and other key operational metrics.
  • Coordinate directly with Cloud Service Providers (CSPs) during infrastructure outages or service disruptions to expedite issue resolution.
  • Continuously assess system reliability, logging, monitoring, and overall architecture, providing recommendations that improve scalability, resiliency, and operational efficiency.
Required Qualifications:
  • TS/SCI security clearance.
  • Strong understanding of Site Reliability Engineering (SRE) principles and best practices.
  • Hands-on experience deploying and managing containerized applications using Kubernetes.
  • Experience administering and troubleshooting multi-cloud environments, including Google Cloud Platform (GCP), Microsoft Azure, and Amazon Web Services (AWS).
  • Experience implementing and maintaining enterprise monitoring, logging, and automated alerting solutions.
  • Proficiency with scripting and automation using languages such as Python, Bash, or similar technologies.
Preferred Qualifications:
  • Passion for building and maintaining highly reliable, mission-critical systems with demanding uptime requirements.
  • Ability to remain composed and methodical while responding to high-priority production incidents.
  • Strong troubleshooting, root cause analysis, and diagnostic skills with a focus on rapid issue resolution.
  • A continuous improvement mindset with an emphasis on automation and eliminating repetitive operational tasks.
  • Experience supporting secure, cloud-native environments within government or defense organizations is a plus.

At IT Veterans LLC, we are committed to providing an environment of mutual respect where equal employment opportunities are available to all applicants and teammates without regard to race, color, religion, sex, pregnancy, national origin, age, physical and mental disability, marital status, sexual orientation, gender identity, gender expression, genetic information, military and veteran status, and any other characteristic protected by applicable law. We believe that diversity and inclusion among our teammates is critical to our success.

#J-18808-Ljbffr