1

Site Reliability Engineer Manager Jobs in California

Site Reliability Engineer

Santa Clara, CA · On-site

$230K - $250K/yr

Forward is transforming how the world's most complex networks are managed and secured. Founded in ... Forward is looking for a Site Reliability Engineer About the Role This is not a "keep the lights on ...

Forward is transforming how the world's most complex networks are managed and secured. Founded in ... Forward is looking for a Site Reliability Engineer About the Role This is not a "keep the lights on ...

Site Reliability Engineer (SRE)

San Diego, CA · On-site

$60.50 - $80.50/hr

Common technologies you'll manage include: Kubernetes (eks), Elasticsearch, Redis, RDS, ELB, and ... managing SRE teams and supporting mission critical applications 3+ years of Hybrid Cloud ...

Senior Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Qualifications BS/MS in Computer Science or Equivalent 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-end project delivery ...

Join us on this journey to redefine resource management-and change lives along the way. The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the reliability ...

next page

Showing results 1-20

Site Reliability Engineer Manager information

See California salary details

$10

$62

$90

How much do site reliability engineer manager jobs pay per hour?

As of Jul 26, 2026, the average hourly pay for site reliability engineer manager in California is $62.91, according to ZipRecruiter salary data. Most workers in this role earn between $54.09 and $71.88 per hour, depending on experience, location, and employer.

Will AI replace SRE jobs?

AI is expected to augment Site Reliability Engineer (SRE) roles by automating routine tasks such as monitoring, incident response, and data analysis, allowing SREs to focus on complex problem-solving and system design. While AI can improve efficiency, it is unlikely to fully replace SREs, as human expertise is essential for managing system architecture, making strategic decisions, and handling unforeseen issues.

What is a Site Reliability Engineer Manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What engineer makes $500,000 a year?

A senior or principal Site Reliability Engineer (SRE) with extensive experience, specialized skills, and often working at large tech companies or in high-cost-of-living areas can earn $500,000 or more annually. Compensation may include base salary, bonuses, and stock options, especially for those in leadership or highly technical roles. Advanced certifications and expertise in cloud platforms, automation, and system architecture are common among top earners in this field.

Is SRE a stressful job?

Site Reliability Engineer (SRE) roles can be stressful due to the high responsibility for system uptime, incident response, and maintaining service reliability. The job often involves working under pressure, handling outages, and balancing automation with manual intervention, but it also offers opportunities for skill development and process improvement. Effective SREs use monitoring tools and incident management practices to manage stress and ensure system stability.

What is the role of site reliability engineer manager?

A Site Reliability Engineer Manager oversees a team responsible for maintaining the reliability, availability, and performance of software systems. They coordinate incident response, implement automation, and ensure system scalability, often using tools like monitoring and alerting platforms. The role requires strong leadership, technical expertise, and knowledge of cloud infrastructure and DevOps practices.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How does a Site Reliability Engineer Manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What are the key skills and qualifications needed to thrive as a Site Reliability Engineer Manager, and why are they important?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.
What are the most commonly searched types of Site Reliability Engineer jobs in California? The most popular types of Site Reliability Engineer jobs in California are:
What cities in California are hiring for Site Reliability Engineer Manager jobs? Cities in California with the most Site Reliability Engineer Manager job openings:
Infographic showing various Site Reliability Engineer Manager job openings in California as of July 2026, with employment types broken down into 90% Full Time, 6% Part Time, and 4% Contract. Highlights an 87% Physical, 5% Hybrid, and 8% Remote job distribution, with an average salary of $130,847 per year, or $62.9 per hour.
Site Reliability Engineer Manager- Hybrid

Site Reliability Engineer Manager- Hybrid

Calance US

Santa Clara, CA

$67.25 - $89.50/hr

Contractor

Medical, Dental, Vision, Life

Posted 17 days ago


Job description

We are hiring Site Reliability Engineer Manager- Hybrid for a Contract To Hire position in santa clara, CA
The Role
You will build and lead the Site Reliability Engineering team, owning the infrastructure that development, validation, and customer-facing deployments run on. This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and the platform services customers use to collaborate on hardware and software deployments.
You are both a people manager and a practicing engineer. You will set technical direction, hire and grow the team, own SLOs for critical systems, and be the senior escalation point when things go wrong. You will work closely with hardware and software development teams to ensure HPC infrastructure meets their workload requirements and partner with the Senior DevOps Lead whose pipelines and automation run on the infrastructure you own.
What You Will Do
Team Leadership & Strategy
Develop and manage a team of 3 5 SRE engineers; establish a culture of operational excellence, ownership, and continuous improvement.
Define the SRE team's technical roadmap: reliability architecture, automation priorities, capacity planning, and on-call model.
Serve as the senior technical escalation for critical incidents guiding cross-team triage, driving RCA, and ensuring systemic fixes rather than point patches.
Translate operational signals and infrastructure health into clear, actionable narratives for engineering leadership and executive stakeholders.
Partner with hardware and software development teams to understand HPC workload requirements and ensure infrastructure capacity, performance, and reliability meet the needs of silicon and software development programs.
24 x7 Infrastructure Reliability & Observability
Own 24 7 reliability across colocation, on-premises lab clusters, cloud, and customer-facing platform services designing for failure domains, progressive delivery, and strict change control at every tier.
Own the full observability stack (metrics, traces, logs) and define SLOs/SLIs across all SRE systems; use AI-driven detection, correlation, and guided remediation to reduce time to detect, respond, and resolve.
Evolve incident and problem management into a data-driven discipline: automated triage workflows, AI/analytics to identify recurring patterns, and every P0/P1 producing a written RCA with tracked systemic fixes.
Lead FinOps and capacity planning: model TCO across cloud vs. on-prem vs. colo, drive workload placement decisions, and anticipate infrastructure needs for new silicon programs and customer deployments.
Own infrastructure for customer collaboration environments where partners deploy and validate hardware and software.
Automation & Infrastructure as Code
Drive IaC-first discipline across the team Terraform, Ansible, and production-quality automation for all infrastructure provisioning and lifecycle management.
Build and mature self-healing infrastructure platforms: host lifecycle automation, fleet auto-remediation, and AIOps-driven alerting that reduce manual intervention across the operational lifecycle.
Documentation & Global Collaboration
Build a documentation culture and scale a follow-the-sun on-call model as we expands globally runbooks, architecture diagrams, and operational playbooks maintained as living artifacts.
Drive POC and POV evaluations for new infrastructure technologies, interconnect fabrics, and platform services relevant to our accelerator roadmap.
What You Will Bring
Required
Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 12+ years in SRE, infrastructure engineering, or production engineering (8 years minimum).
3+ years managing SRE or infrastructure teams hiring, growing, and retaining engineers in a fast-moving environment.
Deep Linux systems expertise: networking (TCP/IP, RDMA, bonding), storage, kernel tuning, and bare-metal operations.
Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking.
IaC fluency: Terraform and Ansible at production scale module design, remote state, environment isolation, and change governance.
Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale.
Full observability stack ownership: Prometheus, Grafana, and/or DataDog SLO definition, alert design, and E2E signal quality.
Strong Python and/or Go production services, not just scripts; automation that touches real infrastructure safely.
Track record of reducing MTTR/MTTD through automation, workflow orchestration, and AIOps tooling.
Executive communication: translating infrastructure health and operational risk into clear narratives for senior leadership.
Demonstrated track record of moving teams from reactive, process-heavy operations to automated, technology-focused models not just managing existing runbooks.
Strongly Preferred
Experience operating customer-facing infrastructure or platform services reliability expectations beyond internal tooling.
Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink setup, troubleshooting, and performance tuning.
HPC job scheduler experience: Slurm, LSF, or equivalent setup, tuning, and integration with infrastructure automation.
Multi-cloud hybrid operations: AWS, Azure, GCP alongside on-prem/colo unified observability and IaC across all tiers.
FinOps: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations for engineering and executive audiences.
ITIL knowledge or equivalent structured incident/problem/change management framework experience.
Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure.
Estimated Pay Range: 90-120/hr