1

Google Site Reliability Engineer Jobs (NOW HIRING)

Site Reliability Engineer - SRE

Atlanta, GA · On-site +1

$54.25 - $72/hr

Role: Site Reliability Engineer * Location: Atlanta, GA OR Dallas OR Austin, TX * Duration ... Strong experience in Google Cloud platform. As a Staff Software Engineer, you will be a core player ...

Site Reliability Engineer - SRE

Atlanta, GA

$54.25 - $72/hr

Role: Site Reliability Engineer * Location: Atlanta, GA OR Dallas OR Austin, TX * Duration ... Strong experience in Google Cloud platform. As a Staff Software Engineer, you will be a core player ...

Site Reliability Engineer

Irondale, AL · On-site

$48.25 - $64/hr

Site Reliability Engineer Site Reliability Engineer (SRE) Hybrid Opportunity | Enterprise Cloud ... Google Cloud Platform (GCP) * Terraform * Prometheus * Grafana * Dynatrace * Azure DevOps (ADO)

Site Reliability Engineer

San Francisco, CA · On-site

$130K - $500K/yr

Deep familiarity with SRE practices as popularized by Google (e.g., error budgets, reliability vs. risk trade-offs, large-scale distributed systems). 5+ years of SRE experience; 15+ years of overall ...

GCP Site Reliability Engineer Interview Mode: candidates local to Parsippany, Nj who can attend an ... Google Cloud Expertise: Design, implement, and manage cloud infrastructure using Google Cloud ...

Site Reliability Engineer

Sterling, VA

$56.50 - $75/hr

The Site Reliability Engineer (SRE) collaboratively works closely with the contract leadership ... Cloud Technologies: Experience with cloud platforms (e.g., AWS, Google Cloud, Azure)

next page

Showing results 1-20

Google Site Reliability Engineer information

See salary details

$10

$63

$91

How much do google site reliability engineer jobs pay per hour?

As of Jul 22, 2026, the average hourly pay for google site reliability engineer in the United States is $63.74, according to ZipRecruiter salary data. Most workers in this role earn between $54.81 and $72.84 per hour, depending on experience, location, and employer.

What are Google Site Reliability Engineers?

Google Site Reliability Engineers (SREs) are specialized engineers responsible for ensuring the reliability, scalability, and performance of Google's infrastructure and services. They bridge the gap between software development and operations by automating tasks, monitoring systems, and responding to incidents. SREs work to minimize downtime, manage large-scale distributed systems, and implement best practices for reliability. Their work involves a combination of software engineering, system administration, and process improvement to support Google's products and users.

How does a Google Site Reliability Engineer typically collaborate with software development teams?

Google Site Reliability Engineers (SREs) work closely with software development teams to ensure that systems are reliable, scalable, and efficient. They participate in design reviews, provide feedback on system architecture with an emphasis on reliability, and help define service level objectives (SLOs) and monitoring standards. SREs often engage in incident response and post-mortem analysis alongside developers, fostering a culture of shared responsibility for uptime and performance. This collaboration helps bridge the gap between development and operations, promoting continuous improvement and rapid innovation.

What is the difference between Google Site Reliability Engineer vs DevOps Engineer?

AspectGoogle Site Reliability EngineerDevOps Engineer
CredentialsTypically requires computer science or engineering degree, experience with coding, systems, and cloud platformsOften has similar technical background, certifications like AWS, Azure, or Linux are common
Work EnvironmentFocuses on maintaining large-scale systems, reliability, and automation at GoogleWorks across development and operations teams to streamline deployment and infrastructure
Industry UsagePrimarily used within Google and similar large tech companiesWidely adopted across various industries for infrastructure and deployment automation

Google Site Reliability Engineers and DevOps Engineers share many skills, including scripting, cloud knowledge, and system management. However, SREs focus more on system reliability and scalability within large-scale environments like Google, while DevOps Engineers emphasize continuous integration, deployment, and collaboration across development and operations teams.

What are the key skills and qualifications needed to thrive as a Google Site Reliability Engineer, and why are they important?

To thrive as a Google Site Reliability Engineer, you need a solid background in computer science, software engineering, and systems administration, often backed by a relevant degree or equivalent experience. Familiarity with programming languages (such as Python, Go, or Java), cloud platforms, configuration management tools, and monitoring systems is essential. Strong problem-solving abilities, effective communication, and a proactive mindset set exceptional SREs apart. These skills are crucial for ensuring the reliability, scalability, and efficiency of large-scale infrastructure and services.
More about Google Site Reliability Engineer jobs
What cities are hiring for Google Site Reliability Engineer jobs? Cities with the most Google Site Reliability Engineer job openings:
What states have the most Google Site Reliability Engineer jobs? States with the most job openings for Google Site Reliability Engineer jobs include:
What job categories do people searching Google Site Reliability Engineer jobs look for? The top searched job categories for Google Site Reliability Engineer jobs are:
Infographic showing various Google Site Reliability Engineer job openings in the United States as of July 2026, with employment types broken down into 100% Full Time. Highlights an 100% In-person job distribution, with an average salary of $132,583 per year, or $63.7 per hour.
Site Reliability Engineer

$142K - $158K/yr

Full-time

Re-posted 18 days ago


Job description

Basic Qualifications
Bachelor's degree in Software Engineering, or related Science, Technology, Engineering or Mathematics field, plus a minimum of 8 years of relevant experience; or Master's degree, plus 6 years relevant experience.
CLEARANCE REQUIREMENTS:: Department of Defense Secret security clearance is required at time of hire. Applicants selected will be subject to a U.S. Government security investigation and must meet eligibility requirements for access to classified information. Due to the nature of work performed within our facilities, U.S. citizenship is required.
Responsibilities for this Position
What You'll Own
  • SLOs and reliability metrics. Define service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions - not just measure uptime.
  • Monitoring and observability. Build and maintain monitoring, logging, and alerting infrastructure for AI services. You will know when something is degrading before users do.
  • Incident response. Establish incident management procedures, lead post-incident reviews, and drive corrective actions. When something breaks, you coordinate the response and ensure it doesn't break the same way again.
  • Operational readiness reviews. Before any AI service goes live, you validate that it meets reliability, security, and operational standards. You are the gate between "it works in dev" and "it's ready for production."
  • Capacity planning and cost monitoring. Track resource consumption, forecast capacity needs, and monitor costs - tokens, compute, storage. You ensure the platform scales without surprises.
  • Toil elimination. Identify and automate repetitive operational tasks. If a human is doing something a script could do, you fix that.
What You Won't Own
  • Application development or AI model building - you ensure what they build is operable, you don't build it
  • Infrastructure provisioning - IT provides the infrastructure; you define what's needed and validate it works
  • Business process decisions or backlog prioritization
What Makes This Role Different
  • AI services have failure modes that traditional applications don't - model drift, token budget exhaustion, prompt injection, upstream data quality degradation. You will build monitoring for problems that most SRE teams have never encountered.
  • You are applying SRE principles from scratch. There is no existing SRE practice to inherit - you will define it for the platform.
  • Your operational readiness reviews directly determine whether AI services go live. You have real authority to say "not ready."
Required Qualifications
  • Bachelor's degree in Computer Science, Software Engineering, or a related field, plus 8 years of experience; or Master's degree plus 6 years of experience
  • Production SRE or DevOps experience - you have owned the reliability of systems that real users depended on, not just built CI/CD pipelines
  • Hands-on experience with monitoring and observability tools - Prometheus, Grafana, Datadog, ELK, CloudWatch, or similar. You have built dashboards and alerts that caught real problems.
  • Strong scripting and automation skills - Python, Bash, infrastructure-as-code (Terraform, CloudFormation, or similar)
  • Experience with containerized environments - Docker, Kubernetes, container orchestration at scale
  • Experience defining and managing SLOs, error budgets, and incident response procedures in production
  • U.S. citizenship required. Department of Defense Secret security clearance is required at time of hire.
Preferred Qualifications
  • Experience with AI/ML production systems - model serving, inference monitoring, token cost tracking, or similar
  • Multi-cloud experience (AWS, Azure, GCP) including cloud-native monitoring and logging services
  • Experience building operational readiness review processes or production launch checklists
  • Familiarity with Google SRE principles - you have read the book and applied the concepts, not just referenced them in interviews
  • Experience in environments where reliability has compliance or safety implications - defense, healthcare, finance, or critical infrastructure
What Sets You Apart
  • You think about failure before you think about features. Your first question about any new system is "how does this break?"
  • You automate yourself out of toil. If you're doing the same thing twice, you write a script.
  • You have said "not ready" to a team that wanted to ship, and you were right.
  • You build monitoring that tells you what's wrong, not just that something is wrong.
  • You write post-incident reviews that actually change how systems are built, not just how incidents are documented.
Details
  • Remote - 100% telework
  • 9/80 schedule
  • Defense industry experience is not required

Salary Note
This estimate represents the typical salary range for this position based on experience and other factors (geographic location, etc.). Actual pay may vary. This job posting will remain open until the position is filled.
Combined Salary Range
USD $142,696.00 - USD $158,303.00 /Yr.
Company Overview
General Dynamics Mission Systems (GDMS) engineers a diverse portfolio of high technology solutions, products and services that enable customers to successfully execute missions across all domains of operation. With a global team of 12,000+ top professionals, we partner with the best in industry to expand the bounds of innovation in the defense and scientific arenas. Given the nature of our work and who we are, we value trust, honesty, alignment and transparency. We offer highly competitive benefits and pride ourselves in being a great place to work with a shared sense of purpose. You will also enjoy a flexible work environment where contributions are recognized and rewarded. If who we are and what we do resonates with you, we invite you to join our high-performance team!
Equal Opportunity Employer / Individuals with Disabilities / Protected Veterans