1

Site Reliability Engineer Manager Jobs in California

SRE

Sunnyvale, CA · On-site

$66.50 - $88.50/hr

Lead and manage SRE projects, ensuring they are delivered on time, within scope, and on budget. • Ensure high customer connect while building processes for all relevant team members to engage with ...

Site Reliability Engineer (SRE)

San Diego, CA · On-site

$60.50 - $80.50/hr

Common technologies you'll manage include: Kubernetes (eks), Elasticsearch, Redis, RDS, ELB, and ... managing SRE teams and supporting mission critical applications 3+ years of Hybrid Cloud ...

Forward is looking for a Site Reliability Engineer About the Role This is not a "keep the lights on ... Experience with network management or observability platforms is a significant plus * Hands-on ...

Site Reliability Engineer

Santa Clara, CA · On-site

$230K - $250K/yr

Forward is looking for a Site Reliability Engineer About the Role This is not a "keep the lights on ... Experience with network management or observability platforms is a significant plus * Hands-on ...

Senior Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Qualifications BS/MS in Computer Science or Equivalent 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-end project delivery ...

Site Reliability Engineer

Santa Clara, CA · On-site

$67.25 - $89.50/hr

Basic understating of Release Management. * Hands-on experience on SNOW. SRE Skillsets ... Expectations from Data Platform team: * Expertise in Message Broker (preferably Rabbit MQ, Kafka)

Common technologies you'll manage include: Kubernetes (eks), Elasticsearch, Redis, RDS, ELB, and ... managing SRE teams and supporting mission critical applications 3+ years of Hybrid Cloud ...

Site Reliability Engineer (SRE)

San Francisco, CA · On-site

$67.25 - $89.25/hr

Methodic is seeking a Site Reliability Engineer (SRE) to focus on the stability and efficiency of ... incident management and a track record of improving systems based on lessons learned. • ...

next page

Showing results 1-20

Site Reliability Engineer Manager information

See California salary details

$10

$62

$90

How much do site reliability engineer manager jobs pay per hour?

As of Aug 25, 2026, the average hourly pay for site reliability engineer manager in California is $62.91, according to ZipRecruiter salary data. Most workers in this role earn between $54.09 and $71.88 per hour, depending on experience, location, and employer.

What is a site reliability engineer manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What are the key skills and qualifications needed to thrive as a site reliability engineer manager?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.

How does a site reliability engineer manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How much do site reliability engineer managers get paid?

Site Reliability Engineer Managers typically earn between $120,000 and $180,000 annually, depending on experience, location, and company size. They often oversee teams responsible for system reliability, incident response, and infrastructure automation, requiring strong leadership and technical skills.

Is a Site Reliability Engineer Manager a stressful job?

A Site Reliability Engineer Manager role can be stressful due to the responsibility of maintaining system uptime, managing incident responses, and ensuring reliability across complex infrastructure. The job often involves working under pressure, handling outages, and coordinating teams, but it also offers opportunities for problem-solving and leadership. Stress levels vary depending on company size, team structure, and workload management skills.

What are the most commonly searched types of Site Reliability Engineer jobs in California?

The most popular types of Site Reliability Engineer jobs in California are:

What cities in California are hiring for Site Reliability Engineer Manager jobs?

Cities in California with the most Site Reliability Engineer Manager job openings:

Infographic showing various Site Reliability Engineer Manager job openings in California as of August 2026, with employment types broken down into 82% Full Time, 15% Part Time, and 3% Contract. Highlights an 79% Physical, 3% Hybrid, and 18% Remote job distribution, with an average salary of $130,847 per year, or $62.9 per hour.

Site Reliability Engineer Manager- Hybrid

Calance US

Santa Clara, CA

$67.25 - $89.50/hr

Contractor

Medical, Dental, Vision, Life

Re-posted 17 days ago


Job description

We are hiring Site Reliability Engineer Manager- Hybrid for a Contract To Hire position in santa clara, CA
The Role
You will build and lead the Site Reliability Engineering team, owning the infrastructure that development, validation, and customer-facing deployments run on. This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and the platform services customers use to collaborate on hardware and software deployments.
You are both a people manager and a practicing engineer. You will set technical direction, hire and grow the team, own SLOs for critical systems, and be the senior escalation point when things go wrong. You will work closely with hardware and software development teams to ensure HPC infrastructure meets their workload requirements and partner with the Senior DevOps Lead whose pipelines and automation run on the infrastructure you own.
What You Will Do
Team Leadership & Strategy
• Develop and manage a team of 3 5 SRE engineers; establish a culture of operational excellence, ownership, and continuous improvement.
• Define the SRE team's technical roadmap: reliability architecture, automation priorities, capacity planning, and on-call model.
• Serve as the senior technical escalation for critical incidents guiding cross-team triage, driving RCA, and ensuring systemic fixes rather than point patches.
• Translate operational signals and infrastructure health into clear, actionable narratives for engineering leadership and executive stakeholders.
• Partner with hardware and software development teams to understand HPC workload requirements and ensure infrastructure capacity, performance, and reliability meet the needs of silicon and software development programs.
24 x7 Infrastructure Reliability & Observability
• Own 24 7 reliability across colocation, on-premises lab clusters, cloud, and customer-facing platform services designing for failure domains, progressive delivery, and strict change control at every tier.
• Own the full observability stack (metrics, traces, logs) and define SLOs/SLIs across all SRE systems; use AI-driven detection, correlation, and guided remediation to reduce time to detect, respond, and resolve.
• Evolve incident and problem management into a data-driven discipline: automated triage workflows, AI/analytics to identify recurring patterns, and every P0/P1 producing a written RCA with tracked systemic fixes.
• Lead FinOps and capacity planning: model TCO across cloud vs. on-prem vs. colo, drive workload placement decisions, and anticipate infrastructure needs for new silicon programs and customer deployments.
• Own infrastructure for customer collaboration environments where partners deploy and validate hardware and software.
Automation & Infrastructure as Code
• Drive IaC-first discipline across the team Terraform, Ansible, and production-quality automation for all infrastructure provisioning and lifecycle management.
• Build and mature self-healing infrastructure platforms: host lifecycle automation, fleet auto-remediation, and AIOps-driven alerting that reduce manual intervention across the operational lifecycle.
Documentation & Global Collaboration
• Build a documentation culture and scale a follow-the-sun on-call model as we expands globally runbooks, architecture diagrams, and operational playbooks maintained as living artifacts.
• Drive POC and POV evaluations for new infrastructure technologies, interconnect fabrics, and platform services relevant to our accelerator roadmap.
What You Will Bring
Required
• Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 12+ years in SRE, infrastructure engineering, or production engineering (8 years minimum).
• 3+ years managing SRE or infrastructure teams hiring, growing, and retaining engineers in a fast-moving environment.
• Deep Linux systems expertise: networking (TCP/IP, RDMA, bonding), storage, kernel tuning, and bare-metal operations.
• Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking.
• IaC fluency: Terraform and Ansible at production scale module design, remote state, environment isolation, and change governance.
• Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale.
• Full observability stack ownership: Prometheus, Grafana, and/or DataDog SLO definition, alert design, and E2E signal quality.
• Strong Python and/or Go production services, not just scripts; automation that touches real infrastructure safely.
• Track record of reducing MTTR/MTTD through automation, workflow orchestration, and AIOps tooling.
• Executive communication: translating infrastructure health and operational risk into clear narratives for senior leadership.
• Demonstrated track record of moving teams from reactive, process-heavy operations to automated, technology-focused models not just managing existing runbooks.
Strongly Preferred
• Experience operating customer-facing infrastructure or platform services reliability expectations beyond internal tooling.
• Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink setup, troubleshooting, and performance tuning.
• HPC job scheduler experience: Slurm, LSF, or equivalent setup, tuning, and integration with infrastructure automation.
• Multi-cloud hybrid operations: AWS, Azure, GCP alongside on-prem/colo unified observability and IaC across all tiers.
• FinOps: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations for engineering and executive audiences.
• ITIL knowledge or equivalent structured incident/problem/change management framework experience.
• Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure.
Estimated Pay Range: 90-120/hr