1

Google Site Reliability Engineer Jobs in Washington

Site Reliability Engineer - Hybrid

Reston, VA ยท On-site

$59.25 - $78.75/hr

Title: Site Reliability Engineer V Location: Reston, VA (Hybrid onsite - 3 days a week from day 1) Assignment duration: 24 months with possibility of extension Interview process: 2 rounds. First ...

Site Reliability Engineer

Chantilly, VA ยท On-site

$62 - $141/hr

R0228604 Site Reliability Engineer The Opportunity Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in ...

Site Reliability Engineer

Washington, DC ยท On-site

$112K - $179K/yr

The SRE will drive automation initiatives, observability improvements, and incident response operations. Site Reliability Engineer responsibilities: * Design and implement automation solutions for ...

The SRE will drive automation initiatives, observability improvements, and incident response operations. Site Reliability Engineer responsibilities: * Design and implement automation solutions for ...

Site Reliability Engineer

Chantilly, VA ยท On-site

$62K - $141K/yr

Site Reliability Engineer The Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network ...

Site Reliability Engineer

Washington, DC ยท On-site

$112K - $179K/yr

The SRE will drive automation initiatives, observability improvements, and incident response operations. Site Reliability Engineer responsibilities: * Design and implement automation solutions for ...

Showing results 41-60

Google Site Reliability Engineer information

See Washington salary details

$12

$72

$104

How much do google site reliability engineer jobs pay per hour?

As of Aug 20, 2026, the average hourly pay for google site reliability engineer in Washington is $72.19, according to ZipRecruiter salary data. Most workers in this role earn between $62.07 and $82.50 per hour, depending on experience, location, and employer.

What is a Google Site Reliability Engineer?

Google Site Reliability Engineers (SREs) are specialized engineers responsible for ensuring the reliability, scalability, and performance of Google's infrastructure and services. They bridge the gap between software development and operations by automating tasks, monitoring systems, and responding to incidents. SREs work to minimize downtime, manage large-scale distributed systems, and implement best practices for reliability. Their work involves a combination of software engineering, system administration, and process improvement to support Google's products and users.

How does a Google Site Reliability Engineer typically collaborate with software development teams?

Google Site Reliability Engineers (SREs) work closely with software development teams to ensure that systems are reliable, scalable, and efficient. They participate in design reviews, provide feedback on system architecture with an emphasis on reliability, and help define service level objectives (SLOs) and monitoring standards. SREs often engage in incident response and post-mortem analysis alongside developers, fostering a culture of shared responsibility for uptime and performance. This collaboration helps bridge the gap between development and operations, promoting continuous improvement and rapid innovation.

What are the key skills and qualifications needed to thrive as a Google Site Reliability Engineer, and why are they important?

To thrive as a Google Site Reliability Engineer, you need a solid background in computer science, software engineering, and systems administration, often backed by a relevant degree or equivalent experience. Familiarity with programming languages (such as Python, Go, or Java), cloud platforms, configuration management tools, and monitoring systems is essential. Strong problem-solving abilities, effective communication, and a proactive mindset set exceptional SREs apart. These skills are crucial for ensuring the reliability, scalability, and efficiency of large-scale infrastructure and services.

What is the difference between Google Site Reliability Engineer vs DevOps Engineer?

AspectGoogle Site Reliability EngineerDevOps Engineer
CredentialsTypically requires computer science or engineering degree, experience with coding, systems, and cloud platformsOften has similar technical background, certifications like AWS, Azure, or Linux are common
Work EnvironmentFocuses on maintaining large-scale systems, reliability, and automation at GoogleWorks across development and operations teams to streamline deployment and infrastructure
Industry UsagePrimarily used within Google and similar large tech companiesWidely adopted across various industries for infrastructure and deployment automation

Google Site Reliability Engineers and DevOps Engineers share many skills, including scripting, cloud knowledge, and system management. However, SREs focus more on system reliability and scalability within large-scale environments like Google, while DevOps Engineers emphasize continuous integration, deployment, and collaboration across development and operations teams.

Is Google Site Reliability Engineer in demand?

The role of a Site Reliability Engineer at Google is highly in demand due to the increasing reliance on cloud services and large-scale infrastructure. SREs with skills in automation, monitoring, and cloud platforms are sought after across the tech industry, reflecting strong job growth in this field.

What are popular job titles related to Google Site Reliability Engineer jobs in Washington?

For Google Site Reliability Engineer jobs in Washington, the most frequently searched job titles are:

What job categories do people searching Google Site Reliability Engineer jobs in Washington look for?

The top searched job categories for Google Site Reliability Engineer jobs in Washington are:

What cities in Washington are hiring for Google Site Reliability Engineer jobs?

Cities in Washington with the most Google Site Reliability Engineer job openings:

Infographic showing various Google Site Reliability Engineer job openings in Washington as of August 2026, with employment types broken down into 1% As Needed, 81% Full Time, 10% Part Time, 2% Temporary, and 6% Contract. Highlights an 89% Physical, 2% Hybrid, and 9% Remote job distribution, with an average salary of $150,163 per year, or $72.2 per hour.

Site Reliability Engineer - Hybrid

Volitiion IIT

Reston, VA โ€ข On-site

$59.25 - $78.75/hr

Other

Re-posted 5 days ago


Job description

Title: Site Reliability Engineer V
Location: Reston, VA (Hybrid onsite - 3 days a week from day 1)
Assignment duration: 24 months with possibility of extension
Interview process: 2 rounds. First round would be a video interview. Second round would be an in-person interview
Manager's call notes
  • This is an SRE role. SRE is under a shared services team within Fannie Mae who works with different application teams. So, multi-tasking is required.
  • In technical terms, we need expertise with AWS ECS, EC2, RDS, RedShift, EMR, Lambda, Route53, Step Functions etc.
  • Programming experience in Java or Python is required. We are not looking for a full fledged developer but someone who can take the code and modify as needed to create some small automations.
  • Exposure to DevOps is required. GitLab, Terraform and Jenkins would be preferred.
  • Experience with Observability using tools such as AWS CloudWatch, Splunk/SignalFX, Dynatrace, and OpenTelemetry would be helpful.
  • If the candidate has experience in release engineering/production support/performance engineering would be a bonus. This is not a show stopped though.
  • AWS, programming and DevOps are must haves.
  • The candidate has to come to the office 3 days a week in Reston, VA.
  • The SRE at Fannie Mae doesn't work 24*7. They get scheduled on a rotation basis. 20% of their job is production support activities. 80% of the time, they work with application teams studying applications, ability to understand the architecture, give suggestions on how the application can be made better, looking into the resiliency patterns and see if the application is resilient enough and suggest new things and work with them. Observability too. Identify gaps and weak points and work with the architecture team to resolve them. Look into the code scans, alarms to see if they are good enough.
  • SRE may not get 100% access to all the applications but the expectation is identifying the gaps/weak points and tell the application team to fix it.
  • On call rotation schedule: One day a week every week.
  • AI/ML: We have certain machine learning projects which the SRE interacts with. So, AI/ML experience is a plus to have.
  • Previous Fannie Mae experience is a plus.
Job Description
Overall years of experience:
8+ years of related experience in their specific area with experience leading teams on projects with similar scope and complexity.
Bachelor's or master's degree in computer science or equivalent.
Certifications: AWS Solutions Architect, Agile Certified Practitioner (ACP), or relevant cloud certifications.
We are seeking a highly skilled and experienced Site Reliability Engineer (SRE) to join our team. The ideal candidate will have a strong background in cloud platforms, DevOps practices, and modern software development frameworks. The SRE will play a critical role in designing, building, and maintaining highly scalable, fault-tolerant, and secure cloud infrastructure while ensuring operational excellence, high availability, and reliability.
Key Responsibilities:
1. Cloud Infrastructure & Automation:
Design, implement, and manage cloud-based infrastructure using platforms like AWS, Azure, or GCP.
Utilize Infrastructure-as-Code (IaC) tools such as Terraform, CloudFormation, and Ansible to automate deployments and configurations.
Create robust automation targeted at anomaly detection, toil reduction, recovery processes, and self-healing mechanisms, and optimize cloud costs.
2. DevSecOps & CI/CD:
Deep understanding of DevSecOps principles and CI/CD pipelines using tools like GitLab, Jenkins, SonarQube, Nexus/Artifactory, and Docker.
Implement security best practices, including IAM roles, RBAC, vulnerability remediation, and SAST/DAST/SCA tools.
3. Observability & Incident Management:
Design and implement monitoring, logging, and distributed tracing solutions using tools like AWS CloudWatch, Splunk/SignalFX, Dynatrace, and OpenTelemetry.
Lead root cause analysis, blameless postmortems, and proactive incident management to minimize MTTR and MTTD.
Define and monitor SLOs, SLIs, and error budgets to ensure system reliability.
4. Microservices & API Management:
Architect and manage microservices, serverless computing, and RESTful APIs.
Ensure fault tolerance and resilience using design patterns like Circuit Breaker, Retry, Timeout, and Bulkhead.
5. Chaos Engineering & Resiliency:
Conduct chaos engineering experiments using tools like AWS FIS and Chaos Toolkit.
Perform resiliency assessments using Resilience Hub and implement self-healing solutions.
6. Database & Application Support:
Manage and optimize database technologies such as PostgreSQL, MongoDB, DynamoDB, Oracle, and Redshift.
Provide production support, including incident response, problem management, and runbook creation. Participate in on-call rotations.
7. Collaboration & Communication:
Collaborate with cross-functional teams to implement shift-left testing practices (BDD, TDD, Unit, Regression).
Create and maintain architecture diagrams, knowledge articles, and disaster recovery plans.
Communicate effectively with stakeholders and demonstrate strong relationship management skills.
Required Skills & Qualifications:
Expertise in cloud platforms (AWS, Azure, or GCP) and container orchestration.
Proficiency in programming/scripting languages such as Python, Java, Node.js, Bash, and PowerShell.
Strong knowledge of database technologies (e.g., PostgreSQL, MongoDB, DynamoDB, Oracle, Redshift).
Experience with DevOps tools (Jenkins, Docker, Nexus/Artifactory) and build tools (Maven, Gradle).
Familiarity with AI/ML integrations, event-driven architectures, and distributed systems.
Expertise in observability, logging, and monitoring tools (AWS CloudWatch, Splunk, Dynatrace, OpenTelemetry).
Strong understanding of security practices, including IAM, RBAC, and vulnerability management.
Experience with chaos engineering, resiliency assessments, and disaster recovery planning.
Proficiency in performance testing tools (JMeter, LoadRunner) and capacity planning.
Excellent verbal and written communication skills, with the ability to collaborate across teams.
Preferred Qualifications:
Experience with AI/ML libraries (e.g., NLTK, Transformers, Spacy, SciPy), Amazon SageMaker, and GenAI tools.
Familiarity with project management tools like JIRA, Confluence, and ServiceNow.
Knowledge of utilities like AWS CLI, POSTMAN, and curl.

4 Reasons to Join Volitiion IIT, Inc.:
1. Our Commitment to You - We offer competitive pay, multi-year projects, and a list of exciting clients.
2. Work-Life Balance - We work hard; we work smart and have quality time for family and "life."
3. Our Mantra - We treat our consultants the way we want to be treated: with integrity, professionalism, and trust.
4. Career Development - We help you meet your career goals and continuously support your efforts to build your skillset.
Check out our Referral Program!
Volitiion IIT Inc will pay you up to $1000 for every qualified professional that you refer and we place. If you see a position posted by Volitiion IIT Inc. and know the perfect person for the job, please send us your referral.

Volitiion IIT Inc. is an Equal Opportunity/Affirmative Action Employer.