The Opportunity We're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10-15 systems and reliability engineers responsible for the availability, performance, and ...
The Opportunity We're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10-15 systems and reliability engineers responsible for the availability, performance, and ...
Job Title- Systems Reliability Team Manager Project Location - Oakland, CA, 94607 hybrid or remote Duration- 12+ months contract Visa- USC/ GC/ GCEAD/ TN Relocation Offered: Yes if submitting remote ...
Quick apply
Job Title- Systems Reliability Team Manager Project Location - Oakland, CA, 94607 hybrid or remote Duration- 12+ months contract Visa- USC/ GC/ GCEAD/ TN Relocation Offered: Yes if submitting remote ...
The Opportunity We're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10-15 systems and reliability engineers responsible for the availability, performance, and ...
The Opportunity We're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10-15 systems and reliability engineers responsible for the availability, performance, and ...
VP, Reliability
Fremont, CA · On-site
$300 - $330/hr
Partner with Product Management and Marketing teams to translate reliability performance into customer-facing benefits, including improved system economics and long-term performance confidence.
VP, Reliability
Fremont, CA · On-site
$300 - $330/hr
Partner with Product Management and Marketing teams to translate reliability performance into customer-facing benefits, including improved system economics and long-term performance confidence.
NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build AI-powered ...
NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build AI-powered ...
The Role We are seeking an exceptional Senior Staff Technical Program Manager (TPM) for Reliability to lead the strategy, execution, and continuous improvement of our most critical Reliability ...
The Role We are seeking an exceptional Senior Staff Technical Program Manager (TPM) for Reliability to lead the strategy, execution, and continuous improvement of our most critical Reliability ...
NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build AI-powered ...
NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build AI-powered ...
Senior Manager, Site Reliability Engineering
Santa Clara, CA · On-site
$248 - $397/hr
NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build AI-powered ...
Senior Manager, Site Reliability Engineering
Santa Clara, CA · On-site
$248 - $397/hr
NVIDIA is seeking a Senior Manager of Site Reliability Engineering to lead and reshape how IT operations function at scale. This role goes beyond traditional service management to build AI-powered ...
The Opportunity We're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10-15 systems and reliability engineers responsible for the availability, performance, and ...
The Opportunity We're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10-15 systems and reliability engineers responsible for the availability, performance, and ...
VP, Reliability
Fremont, CA · On-site
Partner with Product Management and Marketing teams to translate reliability performance into customer-facing benefits, including improved system economics and long-term performance confidence.
VP, Reliability
Fremont, CA · On-site
Partner with Product Management and Marketing teams to translate reliability performance into customer-facing benefits, including improved system economics and long-term performance confidence.
They are seeking a Senior Staff Technical Program Manager for Reliability to lead critical initiatives that enhance the reliability and operational excellence of their multi-cloud infrastructure.
They are seeking a Senior Staff Technical Program Manager for Reliability to lead critical initiatives that enhance the reliability and operational excellence of their multi-cloud infrastructure.
Senior Manager, Cloud Platform & Site Reliability
San Francisco, CA · On-site
$150 - $200/hr
THE ROLE As Senior Manager of Cloud Platform and Site Reliability, you will lead and grow the org responsible for the infrastructure that powers Baseten's machine learning platform. This is a manager ...
Senior Manager, Cloud Platform & Site Reliability
San Francisco, CA · On-site
$150 - $200/hr
THE ROLE As Senior Manager of Cloud Platform and Site Reliability, you will lead and grow the org responsible for the infrastructure that powers Baseten's machine learning platform. This is a manager ...
Responsible for managing and coordinating overall customer quality and reliability requirements specific to foundry manufacturing services, where customer-provided designs are fabricated in Samsung ...
Quick apply
Responsible for managing and coordinating overall customer quality and reliability requirements specific to foundry manufacturing services, where customer-provided designs are fabricated in Samsung ...
Responsible for managing and coordinating overall customer quality and reliability requirements specific to foundry manufacturing services, where customer-provided designs are fabricated in Samsung ...
Responsible for managing and coordinating overall customer quality and reliability requirements specific to foundry manufacturing services, where customer-provided designs are fabricated in Samsung ...
Operations Engineering Manager, Fleet Reliability
$143K - $191K/yr
We are seeking an Operations Manager for the Fleet Reliability Operations team who can help us maintain and improve our high volume of delivery and scale as we 10x the size of our fleet. This ...
Quick apply
Operations Engineering Manager, Fleet Reliability
$143K - $191K/yr
We are seeking an Operations Manager for the Fleet Reliability Operations team who can help us maintain and improve our high volume of delivery and scale as we 10x the size of our fleet. This ...
Manager/Principal Business Unit: Strategy & Growth Work Type: Hybrid Job Location: Oakland Department Overview The System Performance, Reliability and Resiliency Strategy team within the overall ...
Manager/Principal Business Unit: Strategy & Growth Work Type: Hybrid Job Location: Oakland Department Overview The System Performance, Reliability and Resiliency Strategy team within the overall ...
Manager, Reliability, Resiliency, Strategy, and Partnerships
Oakland, CA · On-site
$155K/yr
Manager/Principal Business Unit: Strategy & Growth Work Type: Hybrid Job Location: Oakland Department Overview The System Performance, Reliability and Resiliency Strategy team within the overall ...
Manager, Reliability, Resiliency, Strategy, and Partnerships
Oakland, CA · On-site
$155K/yr
Manager/Principal Business Unit: Strategy & Growth Work Type: Hybrid Job Location: Oakland Department Overview The System Performance, Reliability and Resiliency Strategy team within the overall ...
Role Overview ID.me is seeking a Tech Lead Manager (TLM) to lead and scale our Data Platform Team , with a heavy emphasis on Site Reliability Engineering (SRE) and Data Reliability . This team acts ...
Role Overview ID.me is seeking a Tech Lead Manager (TLM) to lead and scale our Data Platform Team , with a heavy emphasis on Site Reliability Engineering (SRE) and Data Reliability . This team acts ...
Manager/Principal Business Unit: Strategy & Growth Work Type: Hybrid Job Location: Oakland Department Overview The System Performance, Reliability and Resiliency Strategy team within the overall ...
Manager/Principal Business Unit: Strategy & Growth Work Type: Hybrid Job Location: Oakland Department Overview The System Performance, Reliability and Resiliency Strategy team within the overall ...
The Reliability organization evaluates how well every Apple product holds up to the demands of real ... Description As the User Research Engineering Manager, you'll lead a team of researchers with ...
The Reliability organization evaluates how well every Apple product holds up to the demands of real ... Description As the User Research Engineering Manager, you'll lead a team of researchers with ...
Reliability Manager information
See Fremont, CA salary details
$67.9K - $78.5K
8% of jobs
$78.5K - $89.1K
2% of jobs
$89.1K - $99.7K
11% of jobs
$103.4K is the 25th percentile. Wages below this are outliers.
$99.7K - $110.3K
13% of jobs
$110.3K - $120.9K
11% of jobs
The median wage is $127.3K / yr.
$120.9K - $131.5K
10% of jobs
$131.5K - $142.1K
14% of jobs
$149.2K is the 75th percentile. Wages above this are outliers.
$142.1K - $152.7K
11% of jobs
$152.7K - $163.3K
9% of jobs
$163.3K - $173.9K
9% of jobs
$173.9K - $184.5K
4% of jobs
$67.9K
$128.6K
$184.5K
How much do reliability manager jobs pay per year?
What does a reliability manager do?
A Reliability Manager is responsible for ensuring that equipment, processes, and systems operate efficiently and consistently to minimize downtime and maximize performance. They develop and implement reliability strategies, conduct root cause analyses, and oversee preventive and predictive maintenance programs. Their role involves working closely with maintenance teams, engineers, and production staff to improve asset reliability and extend equipment lifespan. Additionally, they analyze failure data, recommend improvements, and help optimize operational costs through reliability-centered maintenance practices.
What are the key skills and qualifications needed to thrive as a reliability manager?
A Reliability Manager needs strong analytical skills, a solid background in engineering or maintenance, and experience with reliability-centered maintenance methodologies. Familiarity with tools like Failure Mode and Effects Analysis (FMEA), Root Cause Analysis (RCA), and certifications such as Certified Reliability Engineer (CRE) are often required. Leadership, problem-solving, and the ability to communicate complex technical information clearly are crucial soft skills for this role. These skills help ensure equipment uptime, optimize maintenance processes, and foster a culture of continuous improvement within the organization.
What are popular job titles related to Reliability Manager jobs in Fremont, CA?
For Reliability Manager jobs in Fremont, CA, the most frequently searched job titles are:
What job categories do people searching Reliability Manager jobs in Fremont, CA look for?
The top searched job categories for Reliability Manager jobs in Fremont, CA are:
What cities near Fremont, CA are hiring for Reliability Manager jobs?
Cities near Fremont, CA with the most Reliability Manager job openings:

Senior Manager, Site Reliability Engineering
Mountain View, CA • On-site
8.2
Based on 92 frontline employees who took The Breakroom Quiz
108th of 246 rated software companies
People enjoy working here
Good employer
Recommended by students
Paid breaks
Recommended by parents
Full-time
Posted 6 days ago
Job description
Intuit's Infrastructure and Site Reliability organization owns the operational backbone that keeps QuickBooks, TurboTax, Credit Karma, and Mailchimp running for hundreds of millions of customers. The Fintech Platform Systems Engineering team builds and operates the AWS-based infrastructure, resiliency tooling, and incident response capability that underpins Intuit's money-movement and fintech services - where availability, data integrity, and trust are non-negotiable.
The OpportunityWe're hiring a Senior Manager, Site Reliability Engineering to lead a hands-on team of 10-15 systems and reliability engineers responsible for the availability, performance, and operational health of Fintech Platform services running in AWS. This leader owns the strategy and execution behind operational excellence: driving toward a 99.999% availability bar, maturing incident management practices, and building self-healing, well-instrumented infrastructure at scale.
This is a player-coach role. You will set technical direction and organizational strategy while staying close to the systems - reviewing designs, joining incident bridges, and coaching engineers through complex production issues. You'll partner closely with software engineering, product, security, and other SRE/infrastructure leaders across Intuit to raise the bar on reliability company-wide.
A defining priority for this role is AI Ops: embedding AI-driven, autonomous operations into how the team runs infrastructure. You will lead the shift from manual, human-triggered response toward self-healing systems that detect, diagnose, and remediate issues autonomously - reducing developer toil, cutting MTTR, and freeing engineering capacity to focus on higher-value work. Done well, this delivers 3x the operational impact of the team today and directly accelerates the pace at which we deliver value to customers.
Responsibilities
ResponsibilitiesOwn end-to-end operational excellence for Fintech Platform services: define and drive the strategy for achieving and sustaining 99.999% availability across customer-facing and internal systems.
Lead, grow, and directly manage a team of 10-15 systems/site reliability engineers - hiring, mentoring, setting goals, and developing the next generation of technical leaders.
Act as a hands-on technical leader: participate in architecture and design reviews, write and review code/IaC where needed, and dive into production systems alongside the team.
- AI Ops: Driving 3x Impact Through Autonomous Operations- Define and execute an AI Ops roadmap that embeds autonomous detection, diagnosis, and remediation into production systems, targeting a 3x improvement
Identify high-toil, repetitive operational workflows and systematically replace them with autonomous agents and automation, freeing engineers to focus on higher-leverage engineering work.
Measure and report on toil reduction, automation coverage, and velocity gains, tying AI Ops investment directly to faster, safer delivery of customer value.
Drive incident management maturity - own the incident command process, lead or oversee response for high-severity (P1/P2) incidents, and ensure rigorous root-cause analysis and blameless postmortems.
Build and scale AWS cloud infrastructure (compute, networking, storage, container orchestration) with a focus on resiliency, auto-remediation, chaos engineering, and multi-AZ/multi-region failover.
Define and report on SLOs/SLIs, error budgets, and availability metrics; use data to prioritize reliability investments and reduce toil through automation.
Partner with software engineering, product management, security, and compliance teams to embed reliability, observability, and operational readiness into the software development lifecycle.
Establish and continuously improve on-call practices, runbooks, alerting, and escalation paths to reduce MTTD/MTTR.
Own capacity planning, cost optimization, and infrastructure roadmap decisions for the systems under your purview.
Represent Infrastructure & SRE in leadership forums, change advisory boards, and executive incident reviews; communicate risk and operational posture clearly to senior stakeholders.
Champion a culture of operational rigor, psychological safety, and continuous improvement across the team.
Qualifications
Qualifications-
8+ years of experience in systems engineering, site reliability engineering, or infrastructure engineering, with 3+ years directly managing engineering teams.
-
Proven, hands-on experience operating production infrastructure in AWS at scale (EC2, EKS/ECS, VPC, RDS/DynamoDB, IAM, CloudWatch, Auto Scaling, and related services).
-
Track record of driving high-availability outcomes (99.9%+ and above) for mission-critical, customer-facing systems, ideally in fintech, payments, or another regulated/high-trust domain.
-
Deep experience with incident management - running incident command, leading postmortems, and building organizational muscle around detection, response, and prevention.
-
Strong technical foundation in distributed systems, networking, containerization/orchestration (Kubernetes), and infrastructure-as-code (Terraform, CloudFormation, or similar).
-
Experience with observability and reliability tooling (e.g., Datadog, Splunk, PagerDuty, Prometheus/Grafana) and building SLO-driven operations.
-
Demonstrated ability to balance hands-on technical depth with people leadership - comfortable reviewing a design doc and coaching a direct report in the same day.
-
Excellent communication skills, with experience presenting risk, status, and strategy to senior/executive leadership.
-
Bachelor's degree in Computer Science, Engineering, or equivalent practical experience.
-
Experience designing or scaling AI Ops / autonomous remediation capabilities (AIOps platforms, ML-based anomaly detection, agentic automation) that measurably reduced toil or MTTR.
-
Experience operating within a regulated fintech, banking, or payments environment (PCI, SOC 2, money movement/ACH systems).
-
Prior experience building or scaling a chaos engineering or resilience testing practice.
-
Familiarity with cost and capacity management for large-scale multi-account AWS environments.
-
Experience leading through major incidents involving cross-functional executive stakeholders.
Footer
Intuit provides a competitive compensation package with a strong pay for performance rewards approach. This position may be eligible for a cash bonus, equity rewards and benefits, in accordance with our applicable plans and programs (see more about our compensation and benefits at Intuit: Careers | Benefits). Pay offered is based on factors such as job-related knowledge, skills, experience, and work location. To drive ongoing fair pay for employees, Intuit conducts regular comparisons across categories of ethnicity and gender.
The expected base pay range for this position is:Mountain View $222,000 - $300,500
Employment Type: Full-Time
About Intuit
Sourced by ZipRecruiter
Industry
Computer and electronic product manufacturing
Company size
5,001 - 10,000 Employees
Headquarters location
Mountain View, CA, US
Year founded
1983