1

Site Reliability Engineer Manager Jobs in California

Site Reliability Engineer

Santa Clara, CA · On-site

$230K - $250K/yr

Forward is looking for a Site Reliability Engineer About the Role This is not a "keep the lights on ... Experience with network management or observability platforms is a significant plus * Hands-on ...

Senior Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Qualifications BS/MS in Computer Science or Equivalent 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-end project delivery ...

Site Reliability Engineer

Palo Alto, CA · On-site

$67 - $89/hr

About the DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-performing production systems. We work closely with ...

Common technologies you'll manage include: Kubernetes (eks), Elasticsearch, Redis, RDS, ELB, and ... managing SRE teams and supporting mission critical applications 3+ years of Hybrid Cloud ...

Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability ...

Site Reliability Engineer (SRE)

San Francisco, CA · On-site

$67.25 - $89.25/hr

Methodic is seeking a Site Reliability Engineer (SRE) to focus on the stability and efficiency of ... incident management and a track record of improving systems based on lessons learned. • ...

Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

... Site Reliability Engineer to own infrastructure and reliability. This role involves designing ... and manage cloud infrastructure efficiently. • Comfortable handling production incidents ...

Site Reliability Engineer

Newport Beach, CA · On-site

$61.25 - $81.50/hr

They are seeking a Site Reliability Engineer to support and maintain the service quality of their ... Obsidian Security provides SaaS and AI security software that detects threats, manages risks, and ...

Site Reliability Engineer (SRE)

Palo Alto, CA · On-site

$67 - $89.25/hr

... to manage Mithril's growing multi-cloud provider footprint. • Write clean, maintainable Python (or Go) to automate repetitive operational tasks -- from provider API reconciliation to automated ...

Site Reliability Engineer

Palo Alto, CA · On-site

$67 - $89.25/hr

They are seeking a Site Reliability Engineer to support and maintain the service quality of their ... Obsidian Security provides SaaS and AI security software that detects threats, manages risks, and ...

Site Reliability Engineer (SRE)

San Francisco, CA · On-site

$67.25 - $89.25/hr

... to manage Mithril's growing multi-cloud provider footprint. • Write clean, maintainable Python (or Go) to automate repetitive operational tasks -- from provider API reconciliation to automated ...

Site Reliability Engineer (SRE)

San Diego, CA · On-site

$142.30 - $263.30/hr

Common technologies you'll manage include: Kubernetes (eks), Elasticsearch, Redis, RDS, ELB, and ... managing SRE teams and supporting mission critical applications * 3+ years of Hybrid Cloud ...

As a Site Reliability Engineer, you will strengthen infrastructure, optimize tooling, deepen ... Configure and improve alerting and routing through incident management workflows. Make sure pages ...

We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You ... Manage datacenter infrastructure (Linux servers, network devices, databases etc.). * Improve ...

Site Reliability Engineer

Mountain View, CA · Hybrid

$67.25 - $89.25/hr

Site Reliability Engineer Onsite- Bay Area, CA Skills Relevant Skills and Experience What You'll Do ... Own and manage our cloud infrastructure (GCP or AWS, on-prem). * Build, maintain, and optimize ...

Showing results 21-40

Site Reliability Engineer Manager information

See California salary details

$10

$62

$90

How much do site reliability engineer manager jobs pay per hour?

As of Aug 9, 2026, the average hourly pay for site reliability engineer manager in California is $62.91, according to ZipRecruiter salary data. Most workers in this role earn between $54.09 and $71.88 per hour, depending on experience, location, and employer.

What is a site reliability engineer manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How does a site reliability engineer manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What are the key skills and qualifications needed to thrive as a site reliability engineer manager?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.
What are the most commonly searched types of Site Reliability Engineer jobs in California? The most popular types of Site Reliability Engineer jobs in California are:
What cities in California are hiring for Site Reliability Engineer Manager jobs? Cities in California with the most Site Reliability Engineer Manager job openings:
Infographic showing various Site Reliability Engineer Manager job openings in California as of July 2026, with employment types broken down into 90% Full Time, 6% Part Time, and 4% Contract. Highlights an 87% Physical, 5% Hybrid, and 8% Remote job distribution, with an average salary of $130,847 per year, or $62.9 per hour.

Site Reliability Engineer

Forward

Santa Clara, CA • On-site

$230K - $250K/yr

Full-time

Posted 25 days ago


Job description

Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically accurate model of the production network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change before it touches production. That founding instinct still defines how we work. We're accurate and evidence-driven, relentless about clarity, and we'd rather be certain than comfortable, building a groundbreaking platform that transforms how teams run and secure networks across every major cloud and vendor environment.
Global leaders like Goldman Sachs, PayPal, S&P Global, IBM, and Dell trust Forward, alongside fast-growing enterprises and government agencies, realizing an average of $14.2 million in annual benefits, according to IDC. Backed by top-tier investors, including A. Capital, Andreessen Horowitz, Goldman Sachs, MSD Partners, Omega Venture Partners, Section 32, and Threshold Ventures, and headquartered in Santa Clara, we're most proud of our team: curious people who'd rather build what doesn't exist than accept how things have always been done.
Forward is looking for a Site Reliability Engineer
About the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward - defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.
If you thrive in environments where you're handed a problem rather than a playbook this role is for you.
What You'll Own
  • Define and drive SRE practices from the ground up - SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use
  • Drive the reliability and operational excellence of the Forward SaaS platform
  • Build and maintain observability infrastructure - logging, metrics, tracing, and alerting - so the team always knows what's happening before customers do
  • Lead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twice
  • Partner with engineering teams to embed reliability thinking into the SDLC - capacity planning, load testing, chaos engineering, and production readiness reviews
  • Help define and build the SRE team as the company scales - this is a foundational hire with a path to leadership

What We're Looking For
  • 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
  • Proven experience building or significantly maturing an SRE function - not just operating within one someone else built
  • Strong fundamentals in networking - TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus
  • Hands-on experience with Kubernetes and container orchestration in production environments
  • Deep proficiency with observability tooling - Prometheus, Grafana, Datadog, Splunk, or similar
  • Strong scripting and automation skills in Python, Bash, or similar
  • Experience with cloud platforms - AWS, GCP, or Azure - including infrastructure as code (Terraform, Ansible, or equivalent)
  • Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements
  • Ability to communicate clearly with both engineering teams and non-technical stakeholders - you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them

Nice to Have
  • Experience supporting enterprise or federal government customers with high availability requirements
  • Experience in a foundational or early SRE hire capacity at a growth stage company

What This Role Is Not
  • A pure ops or NOC role - you are building and engineering, not just monitoring
  • A siloed function - you will be deeply embedded with product and engineering teams
  • A ticket-taker - you will be proactively identifying and solving reliability problems before they become incidents

Why Forward
  • You'll be building something from scratch at a company with real enterprise traction and world-class investors behind it
  • Our customers include some of the most complex network environments on the planet - the reliability bar is high and the work is genuinely interesting
  • People-centric culture built by Stanford Ph.D.s who care deeply about doing things the right way
  • Competitive compensation, equity, and the opportunity to grow into a leadership role as the SRE function scales

The base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location