1

Director Site Reliability Engineering Jobs in Washington

SRE Engineer

Arlington, VA · On-site

$65.75 - $87.25/hr

Responsibilities : • Define, implement, and maintain site reliability engineering practices for mission-critical applications and shared services, with emphasis on uptime, resiliency ...

SRE Engineer

Arlington, VA · On-site

$65.75 - $87.25/hr

Responsibilities : • Define, implement, and maintain site reliability engineering practices for mission-critical applications and shared services, with emphasis on uptime, resiliency ...

... engineering and systems administration practices to ensure the reliability, availability, and ... The SRE will help build resilient systems that scale, automate manual processes, manage fleetwide ...

SRE Engineer

Washington, DC · On-site

$64.50 - $85.75/hr

Reliability Engineering: Champion SRE metrics including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets; design and execute resiliency test plans and support ...

... engineering and systems administration practices to ensure the reliability, availability, and ... The SRE will help build resilient systems that scale, automate manual processes, manage fleetwide ...

Site Reliability Engineer

Sterling, VA

$56.50 - $75/hr

Global Skill Development Council (GSDC) Site Reliability Engineering (SRE) Foundation Certification (CSREF). * AWS Certified SysOps Administrator - Associate. * Google Cloud Certified Professional ...

Site Reliability Engineer

Sterling, VA · On-site

$56.50 - $75/hr

Global Skill Development Council (GSDC) Site Reliability Engineering (SRE) Foundation Certification (CSREF). * AWS Certified SysOps Administrator - Associate. * Google Cloud Certified Professional ...

Site Reliability Engineer (SRE)

Mclean, VA · Remote

$58.25 - $77.50/hr

Site Reliability Engineer (SRE) Remote No sponsorship available. Must be able to obtain a Public ... In this role, you will work closely with engineering, DevOps, cloud, security, and product teams to ...

Site Reliability Engineer (SRE)

Mclean, VA · Remote

$58.25 - $77.50/hr

Site Reliability Engineer (SRE) Remote No sponsorship available. Must be able to obtain a Public ... In this role, you will work closely with engineering, DevOps, cloud, security, and product teams to ...

next page

Showing results 1-20

Director Site Reliability Engineering information

How much do director site reliability engineering get paid?

Director of Site Reliability Engineering typically earns a salary ranging from $150,000 to $250,000 annually, depending on experience, company size, and location. They often oversee teams using tools like Kubernetes and Prometheus and require strong leadership and technical skills.

What is a director site reliability engineering?

A Director of Site Reliability Engineering (SRE) leads teams responsible for ensuring the availability, performance, and scalability of software systems. They define reliability best practices, drive automation, and collaborate with engineering and product teams to improve system resilience. This role requires strong leadership, technical expertise, and a focus on balancing innovation with operational stability.

What are the main challenges faced by a director site reliability engineering, and how can I prepare for them?

A Director of Site Reliability Engineering often encounters challenges such as balancing rapid feature delivery with system stability, managing complex incident responses, and fostering a culture of continuous improvement. Additionally, aligning reliability goals with business objectives and securing cross-functional buy-in can be demanding. To prepare, it is helpful to gain experience in high-scale system management, develop strong leadership and communication abilities, and cultivate a proactive approach to risk management and automation. Staying up to date with the latest SRE practices and building relationships with both engineering and business teams will also support your success in this pivotal role.

What are the key skills and qualifications needed to thrive as a director site reliability engineering?

To thrive as a Director Site Reliability Engineering, you need extensive experience in software engineering, infrastructure management, incident response, and people leadership, often supported by a degree in computer science or a related field. Familiarity with cloud platforms (such as AWS, GCP, or Azure), automation tools (Terraform, Ansible), monitoring systems (Prometheus, Datadog), and relevant certifications like CKA or AWS Solutions Architect is valued. Outstanding communication, stakeholder management, and strategic vision are key soft skills that set leaders apart in this role. These abilities ensure the reliability, scalability, and efficiency of critical systems while effectively guiding and motivating technical teams.

What job categories do people searching Director Site Reliability Engineering jobs in Washington look for? The top searched job categories for Director Site Reliability Engineering jobs in Washington are:
What cities in Washington are hiring for Director Site Reliability Engineering jobs? Cities in Washington with the most Director Site Reliability Engineering job openings:
Infographic showing various Director Site Reliability Engineering job openings in Washington as of August 2026, with employment types broken down into 90% Full Time, and 10% Part Time. Highlights an 70% In-person, 5% Hybrid, and 25% Remote job distribution.

$65.75 - $87.25/hr

Full-time

Re-posted 21 days ago


Job description

Job Summary:
Spatial Front, Inc. is seeking a SRE Engineer to support their growing team. The SRE Engineer will improve the reliability, availability, performance, and operational resilience of mission-critical systems for a federal enterprise program.
Responsibilities:
• Define, implement, and maintain site reliability engineering practices for mission-critical applications and shared services, with emphasis on uptime, resiliency, recoverability, and operational excellence.
• Establish and manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for critical services and environments.
• Implement and maintain monitoring, alerting, and observability solutions for production systems.
• Support production and pre-production operations across development, test, training, staging, and production environments.
• Lead incident response activities, conducting root cause analysis and implementing permanent fixes.
• Support capacity planning, performance analysis, trend monitoring, and scalability planning for enterprise platforms and services.
• Create and maintain runbooks, standard operating procedures, incident playbooks, operational dashboards, and knowledge articles.
• Support high availability, disaster recovery, backup/restore validation, and business continuity activities.
• Develop and implement automation to reduce manual operational toil and improve system reliability.
• Contribute to post-deployment validation, smoke testing, rollback readiness, and environment health checks during releases and maintenance windows.
• Collaborate with teams supporting Oracle/PeopleSoft platforms, integration services, reporting services, and shared enterprise tooling to improve reliability end to end.
• Collaborate with development teams to improve system reliability through design reviews and reliability engineering practices.
• Perform capacity planning and performance optimization for production systems.
• Other duties as assigned.
Qualifications:
Required:
• Bachelor's in Computer Science, Engineering, or related field.
• 5 years software engineering, 3 years site reliability engineering, production support engineering, or platform reliability for enterprise systems, 1 year unix/solaris experience.
• Experience supporting enterprise applications in a high-availability, security-conscious, and compliance-driven environment.
• Experience creating operational documentation, runbooks, and incident response procedures.
• Strong troubleshooting skills across application, middleware, integration, and infrastructure layers.
• Strong verbal and written communication skills, including the ability to work across engineering, security, testing, and program stakeholders.
• Demonstrated expertise in: Site reliability engineering, monitoring, automation, incident response, performance optimization; experienced with UNIX/Solaris.
• Must be a U.S. Citizen.
• Must possess an active Secret security clearance or be able to obtain one.
Preferred:
• DevOps Engineer or equivalent SRE certification.
• Experience supporting environments subject to RMF, STIG, audit, ATO, or similar compliance requirements.
• Experience with Splunk, enterprise monitoring/observability tooling, or similar operational analytics platforms.
• Experience supporting Oracle-based enterprise environments, including Oracle middleware, Oracle Database, or related platform services.
• Experience supporting PeopleSoft or similarly complex ERP / HCM / payroll platforms.
• Exposure to F5, Oracle Data Guard, Oracle GoldenGate, Kafka, or other enterprise integration / traffic / replication technologies.
• Familiarity with scripting and automation using tools such as Shell, Python, or PowerShell.
• Knowledge of DevOps, testing and scanning tools esp. within PeopleSoft environment such as PHIRE, PFT, Tricentis, Palo Alto, CAST etc.
• Experience as an SRE supporting DoD or federal agency programs.
• Familiarity with UNIX/Solaris administration and systems programming.
• Experience with observability platforms such as Prometheus, Grafana, Datadog, or Splunk.
Company:
SFI effectively delivers the right Information Technology solutions and Business Support services using thoughtful analysis, strategic planning and precise execution. Founded in 2008, the company is headquartered in Mc Lean, USA, with a team of 501-1000 employees. The company is currently Late Stage.