1

Site Reliability Engineer Manager Jobs in San Ramon, CA

SRE Engineer

San Jose, CA · On-site

$66.75 - $88.75/hr

Job Title: SRE Engineer Location: San Jose, CA / RTP, NC(Onsite) Job Type: Full Time Must Have Technical/Functional Skills: * SRE, NetApp Storage, Linux Certified, Kubernetes Certified, DevOps, ...

You will drive patch management programs, harden our Cloud infrastructure, and maintain our code ... Site Reliability Engineering, Software Engineering, DevOps, or DevSecOps role. * Demonstrated ...

Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Responsibilities : • Own and execute end-to-end patch management across AWS compute resources ... in a Site Reliability Engineering, Software Engineering, DevOps, or DevSecOps role. • ...

Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Responsibilities : • Own and execute end-to-end patch management across AWS compute resources ... in a Site Reliability Engineering, Software Engineering, DevOps, or DevSecOps role. • ...

Site Reliability Engineer

Sunnyvale, CA · On-site

$145K - $175K/yr

Qualifications (required): * 5-7 years' experience in managing SRE related functions * Expert-level Linux systems administration across complex, production environments (this is a core requirement)

Senior SRE

San Francisco, CA · On-site

$135K - $159K/yr

The Global SRE team is responsible for owning and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Site Reliability Engineer who is ...

We are one of several SRE teams working together to support a platform that serves more than 500,000 customers and manages over 18 million devices worldwide. The team operates with a high degree of ...

Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

As a SRE, you'll be responsible for the reliability, observability, performance, and security of ... Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with ...

Senior SRE

San Francisco, CA · On-site

$127K - $191K/yr

The Global SRE team is responsible for owning and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Site Reliability Engineer who is ...

Site Reliability Engineer (SRE)

Burlingame, CA · On-site

$64.25 - $85.50/hr

They are seeking a Site Reliability Engineer (SRE) to architect and manage the critical ground infrastructure for their satellite constellation, ensuring high availability and seamless integration ...

Senior Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Help manage incident management processes and playbooks . Qualifications * 2+ years of full-time experience in an SRE or similar role * 3+ years of experience working in AWS with EKS and Github (GHA ...

Senior SRE

San Francisco, CA · On-site

$167K - $196K/yr

The Global SRE team is responsible for owning and supporting deployments of global products, and providing first line operational support. We are looking for a Senior Site Reliability Engineer who is ...

Site Reliability Engineer

San Francisco, CA · On-site +1

$152K - $219K/yr

We are one of several SRE teams working together to support a platform that serves more than 500,000 customers and manages over 18 million devices worldwide. The team operates with a high degree of ...

Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Responsibilities : • Own and execute end-to-end patch management across AWS compute resources ... in a Site Reliability Engineering, Software Engineering, DevOps, or DevSecOps role. • ...

Showing results 21-40

Site Reliability Engineer Manager information

See San Ramon, CA salary details

$12

$71

$102

How much do site reliability engineer manager jobs pay per hour?

As of Jul 28, 2026, the average hourly pay for site reliability engineer manager in San Ramon, CA is $71.23, according to ZipRecruiter salary data. Most workers in this role earn between $61.25 and $81.39 per hour, depending on experience, location, and employer.

Will AI replace SRE jobs?

AI is expected to augment Site Reliability Engineer (SRE) roles by automating routine tasks such as monitoring, incident response, and data analysis, allowing SREs to focus on complex problem-solving and system design. While AI can improve efficiency, it is unlikely to fully replace SREs, as human expertise is essential for managing system architecture, making strategic decisions, and handling unforeseen issues.

What is a Site Reliability Engineer Manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What engineer makes $500,000 a year?

A senior or principal Site Reliability Engineer (SRE) with extensive experience, specialized skills, and often working at large tech companies or in high-cost-of-living areas can earn $500,000 or more annually. Compensation may include base salary, bonuses, and stock options, especially for those in leadership or highly technical roles. Advanced certifications and expertise in cloud platforms, automation, and system architecture are common among top earners in this field.

Is SRE a stressful job?

Site Reliability Engineer (SRE) roles can be stressful due to the high responsibility for system uptime, incident response, and maintaining service reliability. The job often involves working under pressure, handling outages, and balancing automation with manual intervention, but it also offers opportunities for skill development and process improvement. Effective SREs use monitoring tools and incident management practices to manage stress and ensure system stability.

What is the role of site reliability engineer manager?

A Site Reliability Engineer Manager oversees a team responsible for maintaining the reliability, availability, and performance of software systems. They coordinate incident response, implement automation, and ensure system scalability, often using tools like monitoring and alerting platforms. The role requires strong leadership, technical expertise, and knowledge of cloud infrastructure and DevOps practices.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How does a Site Reliability Engineer Manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What are the key skills and qualifications needed to thrive as a Site Reliability Engineer Manager, and why are they important?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.
What are the most commonly searched types of Site Reliability Engineer jobs in San Ramon, CA? The most popular types of Site Reliability Engineer jobs in San Ramon, CA are:
What cities near San Ramon, CA are hiring for Site Reliability Engineer Manager jobs? Cities near San Ramon, CA with the most Site Reliability Engineer Manager job openings:
Infographic showing various Site Reliability Engineer Manager job openings in San Ramon, CA as of July 2026, with employment types broken down into 94% Full Time, 3% Part Time, and 3% Contract. Highlights an 88% Physical, 5% Hybrid, and 7% Remote job distribution, with an average salary of $148,164 per year, or $71.2 per hour.

Site Reliability Engineer (SRE)

Thinking Machines Lab

San Francisco, CA • On-site

$350K - $475K/yr

Full-time

Medical, Dental, Vision, PTO

Posted 6 hours ago


Job description

Thinking Machines Lab's mission is to empower humanity through advancing collaborative general intelligence. We're building a future where everyone has access to the knowledge and tools to make AI work for their unique needs and goals.
We are scientists, engineers, and builders who've created some of the most widely used AI products, including ChatGPT and Character.ai, open-weights models like Mistral, as well as popular open source projects like PyTorch, OpenAI Gym, Fairseq, and Segment Anything.
About Tinker
Tinker is our fine-tuning API that empowers researchers and developers to customize frontier AI to their needs - opening access to capabilities that have previously been concentrated in a handful of labs. We manage the infrastructure while allowing Tinkerers full flexibility in training open weights models with their own data, algorithms, and for their own needs. Tinker is rapidly adding new customers, features, and novel use-cases. We're hiring to grow the platform alongside the Tinker community.
About the Role
We're looking for a Site Reliability Engineer to drive the reliability of Tinker end-to-end. You'll work alongside the engineers building the platform and research teams to make every layer of the system more robust and resilient.
What You'll Do
  • Define and own end-to-end reliability, from CI/CD flows to production observability and incident response.
  • Develop appropriate Service Level Objectives for distributed training systems, balancing job completion reliability and scheduling latency with development velocity.
  • Design and implement monitoring and observability across the full training path.
  • Drive incident response for Tinker platform issues, ensuring rapid recovery, thorough incident reviews, and systematic improvements that prevent recurrence.
  • Harden multi-tenant isolation and resource scheduling so that LoRA-based workload co-scheduling maximizes utilization without compromising reliability or data separation
  • Collaborate with security teams to address production vulnerabilities
Skills and Qualifications
Minimum qualifications:
  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.
  • Experience in distributed systems, cloud infrastructure, or site reliability engineering.
  • Proficiency writing software to solve reliability problems, including building tooling and automation.
  • Experience with production incident response, postmortems, and systematic reliability improvement.
  • Strong communication skills and track record of coordination across engineering and research teams.

Preferred qualifications - we encourage you to apply if you meet some but not all of these:
  • Deep experience operating production cloud services at scale (e.g., public cloud platforms, internal cloud services)
  • Background in distributed training frameworks and how infrastructure failures surface in training behavior.
  • Track record building checkpoint and recovery systems for long-running distributed jobs.
  • Expertise in Kubernetes at scale: deploying, operating, debugging, and tuning clusters handling heterogeneous GPU workloads.
Logistics
  • Location: This role is based in San Francisco, California.
  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.
Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.