1

Director Site Reliability Engineering Jobs in California

Site Reliability Engineer

San Francisco, CA · Remote

$67.25 - $89.25/hr

Director, Site Reliability Location: Remote (US) Department: Cloud Platform Engineering / SRE/Reliability Position summary The Site Reliability Engineer (SRE) owns reliability, observability, and ...

Site Reliability Engineering

Los Angeles, CA · On-site

$61.50 - $81.50/hr

Site Reliability Engineering (SRE) Location: Los Angeles, CA Remote position Fulltime position JD * Site Reliability Engineer * Experience in Cloud platforms (AWS, Azure, Google Cloud) and hybrid ...

Serve as subject matter expert in an SRE mindset, best practices, and cloud-native principles * Scale systems sustainably through automation to improve reliability and velocity * Assist with all ...

Site Reliability Engineering

San Francisco, CA

$67.25 - $89.25/hr

Serve as subject matter expert in an SRE mindset, best practices, and cloud-native principles * Scale systems sustainably through automation to improve reliability and velocity * Assist with all ...

Site Reliability Engineering

Los Angeles, CA

$61.50 - $81.50/hr

Serve as subject matter expert in an SRE mindset, best practices, and cloud-native principles * Scale systems sustainably through automation to improve reliability and velocity * Assist with all ...

Manager, Site Reliability Engineering

San Francisco, CA · Hybrid

$67.25 - $89.25/hr

Manager, Site Reliability Engineering San Francisco, California Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted ...

Manager, Site Reliability Engineering

San Francisco, CA · Hybrid

$67.25 - $89.25/hr

Manager, Site Reliability Engineering San Francisco, California Secure Every Identity, from AI to Human Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted ...

next page

Showing results 1-20

Director Site Reliability Engineering information

See California salary details

$10

$62

$90

How much do director site reliability engineering jobs pay per hour?

As of Aug 8, 2026, the average hourly pay for director site reliability engineering in California is $62.91, according to ZipRecruiter salary data. Most workers in this role earn between $54.09 and $71.88 per hour, depending on experience, location, and employer.

How much do director site reliability engineering get paid?

Director of Site Reliability Engineering typically earns a salary ranging from $150,000 to $250,000 annually, depending on experience, company size, and location. They often oversee teams using tools like Kubernetes and Prometheus and require strong leadership and technical skills.

What is a director site reliability engineering?

A Director of Site Reliability Engineering (SRE) leads teams responsible for ensuring the availability, performance, and scalability of software systems. They define reliability best practices, drive automation, and collaborate with engineering and product teams to improve system resilience. This role requires strong leadership, technical expertise, and a focus on balancing innovation with operational stability.

What are the main challenges faced by a director site reliability engineering, and how can I prepare for them?

A Director of Site Reliability Engineering often encounters challenges such as balancing rapid feature delivery with system stability, managing complex incident responses, and fostering a culture of continuous improvement. Additionally, aligning reliability goals with business objectives and securing cross-functional buy-in can be demanding. To prepare, it is helpful to gain experience in high-scale system management, develop strong leadership and communication abilities, and cultivate a proactive approach to risk management and automation. Staying up to date with the latest SRE practices and building relationships with both engineering and business teams will also support your success in this pivotal role.

What are the key skills and qualifications needed to thrive as a director site reliability engineering?

To thrive as a Director Site Reliability Engineering, you need extensive experience in software engineering, infrastructure management, incident response, and people leadership, often supported by a degree in computer science or a related field. Familiarity with cloud platforms (such as AWS, GCP, or Azure), automation tools (Terraform, Ansible), monitoring systems (Prometheus, Datadog), and relevant certifications like CKA or AWS Solutions Architect is valued. Outstanding communication, stakeholder management, and strategic vision are key soft skills that set leaders apart in this role. These abilities ensure the reliability, scalability, and efficiency of critical systems while effectively guiding and motivating technical teams.

What job categories do people searching Director Site Reliability Engineering jobs in California look for? The top searched job categories for Director Site Reliability Engineering jobs in California are:
What cities in California are hiring for Director Site Reliability Engineering jobs? Cities in California with the most Director Site Reliability Engineering job openings:
Infographic showing various Director Site Reliability Engineering job openings in California as of July 2026, with employment types broken down into 87% Full Time, and 13% Part Time. Highlights an 76% In-person, 4% Hybrid, and 20% Remote job distribution, with an average salary of $130,847 per year, or $62.9 per hour.

Director, Site Reliability Engineering

Salesforce, Inc.

San Francisco, CA

$67.25 - $89.25/hr

Full-time

Medical, Dental, Vision, Life, Retirement

Posted 3 days ago

New


Salesforce rating

8.1

Company rating: 8.1 out of 10

Based on 58 frontline employees who took The Breakroom Quiz

112th of 242 rated software companies


Job description

To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.

Job Category

Software Engineering

Job Details

About Salesforce

Salesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn't a buzzword - it's a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.

Ready to level-up your career at the company leading workforce transformation in the agentic era? You're in the right place! Agentforce is the future of AI, and you are the future of Salesforce.

Job Title: Director, Site Reliability Engineering
Location: New York, NY; San Francisco, CA; Dallas, TX

About the Role

We are looking for a Director of Site Reliability Engineering to spearhead the evolution of our reliability, observability, and operational engineering capabilities.

In this role, you will transform our SRE function-moving our engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. Partnering closely across Application Engineering, Platform, Architecture, Security, Infrastructure, and Product, you will ensure our services are resilient, observable, scalable, and production-ready long before they launch.

As an impactful people leader with sharp technical judgment, you will directly manage and empower a core team of ~6 engineers while driving cross-functional alignment across a complex organization. You won't just run existing playbooks; you will define the strategy, tooling, automation, and culture needed to mentor your team and scale system reliability enterprise-wide.

Key Responsibilities

SRE Strategy and Leadership

  • Define and execute the long-term strategy and roadmap for Site Reliability Engineering.

  • Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures.

  • Build and develop a high-performing team of site reliability and operations engineers.

  • Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices.

  • Translate business priorities and customer impact into clear reliability investments and engineering outcomes.

  • Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs.

Reliability Engineering

  • Establish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services.

  • Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems.

  • Define what it means for a service to be operationally and observably ready for production.

  • Develop readiness reviews and certification practices for high-impact services and launches.

  • Drive improvements in system availability, performance, resiliency, and recovery.

  • Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment.

Observability

  • Define an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry.

  • Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards.

  • Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities.

  • Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure.

  • Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes.

  • Establish governance and measurement to assess adoption and effectiveness of observability standards.

Incident Management and Operational Excellence

  • Improve incident detection, response, mitigation, communication, and learning.

  • Lead the transition from manual and reactive operations toward automated detection, diagnosis, remediation, and incident creation.

  • Reduce mean time to detect, acknowledge, mitigate, and recover.

  • Improve on-call practices, escalation paths, runbooks, and operational ownership.

  • Establish blameless post-incident review practices that produce measurable engineering improvements.

  • Identify recurring sources of operational toil and create plans to eliminate or automate them.

  • Partner with engineering leaders to ensure actions from incidents are prioritized and completed.

Automation and AI-Enabled Operations

  • Develop a roadmap for intelligent operations, including anomaly detection, event correlation, automated triage, assisted root-cause analysis, and remediation.

  • Evaluate opportunities to use agents and AI-assisted workflows across observability, incident response, capacity planning, and operational support.

  • Build automation that reduces cognitive load and improves the speed and consistency of operational decisions.

  • Ensure automation is safe, measurable, auditable, and designed with appropriate human oversight.

  • Promote platform and self-service approaches that allow product teams to adopt reliability practices with minimal friction.

Cross-Functional Partnership

  • Partner with engineering, DevOps, and business stakeholders.

  • Influence teams that do not directly report into SRE and build shared accountability for production outcomes.

  • Create clear service ownership models and operational expectations across teams.

  • Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination.

  • Communicate reliability posture, risks, trends, and investments to executive and technical audiences.

Measures of Success

  • Improved availability and reliability of critical services.

  • Reduced time to detect, diagnose, mitigate, and recover from incidents.

  • Increased percentage of services meeting observability and production-readiness standards.

  • Reduced alert noise, operational toil, and manual incident-management activity.

  • Increased adoption of service-level objectives and measurable reliability practices.

  • Improved quality and completion rate of post-incident corrective actions.

  • Increased automation across detection, triage, remediation, and operational workflows.

  • Stronger ownership of production reliability across engineering teams.

Leadership Attributes
  • Engineering-first and automation-oriented.

  • Comfortable challenging legacy operating models and assumptions.

  • Able to move between technical detail and executive-level strategy.

  • Outcome-focused, pragmatic, and data-driven.

  • Builds trust through clarity, accountability, and strong partnership.

  • Develops leaders and creates an inclusive, high-performance engineering culture.

  • Treats incidents as opportunities to improve systems rather than assign blame.

  • Brings urgency to operational risks while maintaining focus on sustainable solutions.

Minimum Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field; Master's degree or MBA preferred.

  • 10+ years of progressive engineering experience, including 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams

  • Proven experience building or transforming a reliability or operational engineering organization.

  • Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery.

  • Experience establishing observability, incident-management, service-level objective, and production-readiness practices.

  • Demonstrated ability to improve reliability through engineering and automation rather than process alone.

  • Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems.

  • Strong understanding of modern telemetry, including metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring.

  • Demonstrated success driving alignment and building consensus across cross-functional engineering teams and executive stakeholders

  • Ability to balance immediate operational needs with long-term engineering transformation.

  • Strong written, verbal, and executive communication skills.

Preferred Qualifications
  • Strong experience operating large-scale systems in AWS or another major cloud environment.

  • Proven track record with observability platforms such as New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry.

  • Demonstrated experience implementing OpenTelemetry or common instrumentation standards.

  • Verified proficiency building internal developer platforms, paved roads, or self-service reliability capabilities.

  • Experience applying AI, machine learning, or agent-based automation to operational workflows.

  • Seasoned capability with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering.

  • Software engineering experience and the ability to engage deeply in architecture and design discussions.

  • Solid background in supporting high-profile launches, events, or systems with significant customer and business impact.

*LI-Y

Unleash Your Potential

When you join Salesforce, you'll be limitless in all areas of your life. Our benefits and resources support you to find balance andbe your best, and our AI agents accelerate your impact so you cando your best. Together, we'll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Apply today to not only shape the future - but to redefine what's possible - for yourself, for AI, and the world.

Accommodations

If you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form.

Please note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates' resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.

Posting Statement

Salesforce is an equal opportunity employer and maintains a policy of non-discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that's inclusive, and free from discrimination. Know your rights: workplace discrimination is illegal. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications - without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law. This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between. Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education.

In the United States, compensation offered will be determined by factors such as location, job level, job-related knowledge, skills, and experience. Certain roles may be eligible for incentive compensation, equity, and benefits. Salesforce offers a variety of benefits to help you live well including: time off programs, medical, dental, vision, mental health support, paid parental leave, life and disability insurance, 401(k), and an employee stock purchasing program. More details about company benefits can be found at the following link: https://www.salesforcebenefits.com.Pursuant to the San Francisco Fair Chance Ordinance and the Los Angeles Fair Chance Initiative for Hiring, Salesforce will consider for employment qualified applicants with arrest and conviction records.At Salesforce, we believe in equitable compensation practices that reflect the dynamic nature of labor markets across various regions. The typical base salary range for this position is $197,300 - $313,700 annually. In select cities within the San Francisco and New York City metropolitan area, the base salary range for this role is $237,700 - $344,700 annually. The range represents base salary only, and does not include company bonus, incentive for sales roles, equity or benefits, as applicable.

What Salesforce employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom