1

Sre Manager Jobs in California (NOW HIRING)

Site Reliability Engineer

San Francisco, CA · Remote

$67.25 - $89.25/hr

Site Reliability Engineer Platform and software · shared across customers Reports to: Director ... Manage on-call rotation, escalation paths, and incident-management tooling * Coordinate cross ...

Site Reliability Engineer (SRE)

San Francisco, CA · On-site

$67.25 - $89.25/hr

As a Site Reliability Engineer (SRE), you will be responsible for ensuring the reliability ... management, debugging, and root cause analysis. • Proficiency in scripting (Bash, Python, or Go ...

Site Reliability Engineer (SRE)

San Diego, CA · On-site

$60.50 - $80.50/hr

Common technologies you'll manage include: Kubernetes (eks), Elasticsearch, Redis, RDS, ELB, and ... managing SRE teams and supporting mission critical applications 3+ years of Hybrid Cloud ...

Forward is looking for a Site Reliability Engineer About the Role This is not a "keep the lights on ... Experience with network management or observability platforms is a significant plus * Hands-on ...

Senior Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Qualifications BS/MS in Computer Science or Equivalent 6+ years of experience in Site Reliability Engineering, managing infrastructure and services at scale History of end-to-end project delivery ...

Site Reliability Engineer

Santa Clara, CA · On-site

$230K - $250K/yr

Forward is looking for a Site Reliability Engineer About the Role This is not a "keep the lights on ... Experience with network management or observability platforms is a significant plus * Hands-on ...

next page

Showing results 1-20

Sre Manager information

See California salary details

$61.2K

$115.9K

$166.3K

How much do sre manager jobs pay per year?

As of Aug 10, 2026, the average yearly pay for sre manager in California is $115,949.00, according to ZipRecruiter salary data. Most workers in this role earn between $93,300.00 and $138,200.00 per year, depending on experience, location, and employer.

What is an SRE manager?

SRE Managers are leaders responsible for overseeing Site Reliability Engineering (SRE) teams. They ensure the reliability, scalability, and performance of software systems by guiding engineers in implementing best practices, automation, and monitoring processes. SRE Managers collaborate closely with development and operations teams to balance feature development with system stability. Their role also includes mentoring SREs, managing incident response, and driving improvements in system reliability and operational efficiency.

What are the key skills and qualifications needed to thrive as an SRE manager, and why are they important?

To thrive as an SRE Manager, you need a deep understanding of site reliability engineering principles, strong experience with systems architecture, and a background in computer science or a related field. Familiarity with tools such as Kubernetes, Prometheus, cloud platforms, and CI/CD pipelines, as well as certifications like AWS Certified Solutions Architect, are commonly expected. Leadership, effective communication, and problem-solving skills are crucial for driving team performance and collaborating across departments. These skills ensure high system reliability, efficient incident management, and a culture of continuous improvement within technical organizations.

What are some common challenges SRE managers face when leading Site Reliability Engineering teams?

SRE Managers often encounter challenges balancing reliability with rapid development, ensuring their teams have the right mix of software engineering and operations skills. They must also foster a culture of continuous improvement while managing on-call rotations and incident response without causing burnout. Additionally, collaborating effectively with development and product teams to set realistic service level objectives (SLOs) and drive adoption of SRE best practices can require strong communication and negotiation skills.

Is SRE manager a good career path?

An SRE (Site Reliability Engineering) manager oversees teams responsible for maintaining system reliability, scalability, and performance. It is a high-demand role that requires strong technical skills, leadership, and knowledge of tools like monitoring systems and automation. This career path offers growth opportunities in cloud environments and DevOps practices, making it a solid choice for those interested in infrastructure and operations management.

What is the difference between Sre Manager vs DevOps Engineer?

AspectSre ManagerDevOps Engineer
CredentialsTypically requires a Bachelor's/Master's in CS or related field, with certifications like AWS, Google Cloud, or KubernetesSimilar credentials, often with cloud certifications and scripting skills
Work EnvironmentLeads teams, manages incident response, and oversees reliability strategiesFocuses on automation, CI/CD pipelines, and infrastructure deployment
Industry UsageCommon in large tech companies, financial services, and cloud providersWidely used across startups, tech firms, and enterprises adopting DevOps practices

The Sre Manager and DevOps Engineer roles share overlapping skills in cloud computing, automation, and infrastructure management. While the Sre Manager oversees reliability and team coordination, the DevOps Engineer focuses on implementing automation tools and deployment pipelines. Both roles are crucial for modern IT operations, but the Sre Manager typically has a broader leadership responsibility, whereas the DevOps Engineer is more hands-on with technical implementation.

What are the most commonly searched types of Sre jobs in California? The most popular types of Sre jobs in California are:
What job categories do people searching Sre Manager jobs in California look for? The top searched job categories for Sre Manager jobs in California are:
What cities in California are hiring for Sre Manager jobs? Cities in California with the most Sre Manager job openings:
Infographic showing various Sre Manager job openings in California as of August 2026, with employment types broken down into 1% As Needed, 83% Full Time, 13% Part Time, 2% Temporary, and 1% Contract. Highlights an 93% Physical, 3% Hybrid, and 4% Remote job distribution, with an average salary of $115,949 per year, or $55.7 per hour.

Site Reliability Engineer Manager- Hybrid

Calance US

Santa Clara, CA

$67.25 - $89.50/hr

Contractor

Medical, Dental, Vision, Life

Re-posted 2 days ago


Job description

We are hiring Site Reliability Engineer Manager- Hybrid for a Contract To Hire position in santa clara, CA
The Role
You will build and lead the Site Reliability Engineering team, owning the infrastructure that development, validation, and customer-facing deployments run on. This spans colocation facilities, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and the platform services customers use to collaborate on hardware and software deployments.
You are both a people manager and a practicing engineer. You will set technical direction, hire and grow the team, own SLOs for critical systems, and be the senior escalation point when things go wrong. You will work closely with hardware and software development teams to ensure HPC infrastructure meets their workload requirements and partner with the Senior DevOps Lead whose pipelines and automation run on the infrastructure you own.
What You Will Do
Team Leadership & Strategy
Develop and manage a team of 3 5 SRE engineers; establish a culture of operational excellence, ownership, and continuous improvement.
Define the SRE team's technical roadmap: reliability architecture, automation priorities, capacity planning, and on-call model.
Serve as the senior technical escalation for critical incidents guiding cross-team triage, driving RCA, and ensuring systemic fixes rather than point patches.
Translate operational signals and infrastructure health into clear, actionable narratives for engineering leadership and executive stakeholders.
Partner with hardware and software development teams to understand HPC workload requirements and ensure infrastructure capacity, performance, and reliability meet the needs of silicon and software development programs.
24 x7 Infrastructure Reliability & Observability
Own 24 7 reliability across colocation, on-premises lab clusters, cloud, and customer-facing platform services designing for failure domains, progressive delivery, and strict change control at every tier.
Own the full observability stack (metrics, traces, logs) and define SLOs/SLIs across all SRE systems; use AI-driven detection, correlation, and guided remediation to reduce time to detect, respond, and resolve.
Evolve incident and problem management into a data-driven discipline: automated triage workflows, AI/analytics to identify recurring patterns, and every P0/P1 producing a written RCA with tracked systemic fixes.
Lead FinOps and capacity planning: model TCO across cloud vs. on-prem vs. colo, drive workload placement decisions, and anticipate infrastructure needs for new silicon programs and customer deployments.
Own infrastructure for customer collaboration environments where partners deploy and validate hardware and software.
Automation & Infrastructure as Code
Drive IaC-first discipline across the team Terraform, Ansible, and production-quality automation for all infrastructure provisioning and lifecycle management.
Build and mature self-healing infrastructure platforms: host lifecycle automation, fleet auto-remediation, and AIOps-driven alerting that reduce manual intervention across the operational lifecycle.
Documentation & Global Collaboration
Build a documentation culture and scale a follow-the-sun on-call model as we expands globally runbooks, architecture diagrams, and operational playbooks maintained as living artifacts.
Drive POC and POV evaluations for new infrastructure technologies, interconnect fabrics, and platform services relevant to our accelerator roadmap.
What You Will Bring
Required
Bachelor's or Master's in Computer Science, Electrical Engineering, or related field; 12+ years in SRE, infrastructure engineering, or production engineering (8 years minimum).
3+ years managing SRE or infrastructure teams hiring, growing, and retaining engineers in a fast-moving environment.
Deep Linux systems expertise: networking (TCP/IP, RDMA, bonding), storage, kernel tuning, and bare-metal operations.
Proven experience operating colocation and on-premises hardware at scale: server lifecycle, power and cooling awareness, rack-level networking.
IaC fluency: Terraform and Ansible at production scale module design, remote state, environment isolation, and change governance.
Kubernetes cluster operations: lifecycle management, workload reliability, storage, and RBAC at scale.
Full observability stack ownership: Prometheus, Grafana, and/or DataDog SLO definition, alert design, and E2E signal quality.
Strong Python and/or Go production services, not just scripts; automation that touches real infrastructure safely.
Track record of reducing MTTR/MTTD through automation, workflow orchestration, and AIOps tooling.
Executive communication: translating infrastructure health and operational risk into clear narratives for senior leadership.
Demonstrated track record of moving teams from reactive, process-heavy operations to automated, technology-focused models not just managing existing runbooks.
Strongly Preferred
Experience operating customer-facing infrastructure or platform services reliability expectations beyond internal tooling.
Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink setup, troubleshooting, and performance tuning.
HPC job scheduler experience: Slurm, LSF, or equivalent setup, tuning, and integration with infrastructure automation.
Multi-cloud hybrid operations: AWS, Azure, GCP alongside on-prem/colo unified observability and IaC across all tiers.
FinOps: cloud spend attribution, TCO modeling across cloud vs. on-prem vs. colo, and translating cost data into workload placement recommendations for engineering and executive audiences.
ITIL knowledge or equivalent structured incident/problem/change management framework experience.
Published technical writing, conference talks, or open-source contributions in reliability, observability, or HPC infrastructure.
Estimated Pay Range: 90-120/hr