1

Director Site Reliability Engineering Jobs in Utah

$57.75 - $76.75/hr

Site Reliability Engineer (SRE) Department: Technology Location: Manila Reporting To: Head of Infra ... You will collaborate with engineering, DevOps, and client success teams to operationalize ...

Site Reliability Engineer

Saint George, UT · On-site

$50.75 - $67.50/hr

Shares best practices - Shares understanding of Site Reliability Engineering culture across organization; shares knowledge of best practices, approaches, documentation, and code with team members and ...

Site Reliability Engineer

Saint George, UT · On-site

$53.75 - $71.50/hr

Shares best practices - Shares understanding of Site Reliability Engineering culture across organization; shares knowledge of best practices, approaches, documentation, and code with team members and ...

Site Reliability Engineer II

Lehi, UT · On-site

$53.50 - $71/hr

ABOUT THIS POSITION We are looking for a talented and driven Site Reliability Engineering (SRE) to support our engineering team, which manages the infrastructure and services that power our Waystar ...

Sr. Site Reliability Engineer

Salt Lake City, UT · On-site

$55.25 - $73.25/hr

Salt Lake City, UT As a Senior Site Reliability Engineer, you will help define the future of ... You'll leverage software engineering, automation, and cloud-native technologies to build and ...

Staff SRE

Pleasant Grove, UT · On-site

$51.50 - $68.25/hr

The Staff Site Reliability Engineer will architect and implement enterprise-scale infrastructure ... leadership across engineering teams to ensure exceptional system reliability and performance.

Site Reliability Engineer

Draper, UT · On-site

$53.25 - $70.75/hr

... hardware engineering teams to define Service Level Indicators (SLIs), Service Level Objectives ... an SRE, DevOps, or Infrastructure Engineering role supporting production cloud environments • ...

Staff SRE

Pleasant Grove, UT · On-site

$51.50 - $68.25/hr

In this high-impact staff-level role, the Staff Site Reliability Engineer will architect and ... with engineering and security leadership. • Lead performance optimization initiatives ...

Senior Site Reliability Engineer

Lehi, UT

$53.50 - $71/hr

This role bridges software engineering and cloud operations, ensuring mission-critical systems remain highly available and resilient. By integrating reliability early, the SRE fosters a culture of ...

Security Site Reliability Engineer

Pleasant Grove, UT · On-site

$51.50 - $68.25/hr

Security Site Reliability Engineer Join Us at Pura-Reimagining Fragrance for the Future At Pura, we ... You'll work alongside our AppSec engineer and Security Director to protect the infrastructure that ...

Senior Site Reliability Engineer

Lehi, UT · On-site

$53.50 - $71/hr

This role bridges software engineering and cloud operations, ensuring mission-critical systems remain highly available and resilient. By integrating reliability early, the SRE fosters a culture of ...

next page

Showing results 1-20

Director Site Reliability Engineering information

See Utah salary details

$9

$58

$83

How much do director site reliability engineering jobs pay per hour?

As of Jul 30, 2026, the average hourly pay for director site reliability engineering in Utah is $58.03, according to ZipRecruiter salary data. Most workers in this role earn between $49.90 and $66.30 per hour, depending on experience, location, and employer.

What is a Director Site Reliability Engineering job?

A Director of Site Reliability Engineering (SRE) leads teams responsible for ensuring the availability, performance, and scalability of software systems. They define reliability best practices, drive automation, and collaborate with engineering and product teams to improve system resilience. This role requires strong leadership, technical expertise, and a focus on balancing innovation with operational stability.

What are the main challenges faced by a Director of Site Reliability Engineering, and how can I prepare for them?

A Director of Site Reliability Engineering often encounters challenges such as balancing rapid feature delivery with system stability, managing complex incident responses, and fostering a culture of continuous improvement. Additionally, aligning reliability goals with business objectives and securing cross-functional buy-in can be demanding. To prepare, it is helpful to gain experience in high-scale system management, develop strong leadership and communication abilities, and cultivate a proactive approach to risk management and automation. Staying up to date with the latest SRE practices and building relationships with both engineering and business teams will also support your success in this pivotal role.

What are the key skills and qualifications needed to thrive in the Director Site Reliability Engineering position, and why are they important?

To thrive as a Director Site Reliability Engineering, you need extensive experience in software engineering, infrastructure management, incident response, and people leadership, often supported by a degree in computer science or a related field. Familiarity with cloud platforms (such as AWS, GCP, or Azure), automation tools (Terraform, Ansible), monitoring systems (Prometheus, Datadog), and relevant certifications like CKA or AWS Solutions Architect is valued. Outstanding communication, stakeholder management, and strategic vision are key soft skills that set leaders apart in this role. These abilities ensure the reliability, scalability, and efficiency of critical systems while effectively guiding and motivating technical teams.

What job categories do people searching Director Site Reliability Engineering jobs in Utah look for? The top searched job categories for Director Site Reliability Engineering jobs in Utah are:
What cities in Utah are hiring for Director Site Reliability Engineering jobs? Cities in Utah with the most Director Site Reliability Engineering job openings:
Infographic showing various Director Site Reliability Engineering job openings in Utah as of July 2026, with employment types broken down into 1% As Needed, 86% Full Time, 10% Part Time, 1% Temporary, and 2% Contract. Highlights an 92% Physical, 3% Hybrid, and 5% Remote job distribution, with an average salary of $120,700 per year, or $58 per hour.

Manager, Site Reliability Engineering

Octanner

Salt Lake City, UT • On-site

Full-time

Posted 9 days ago


Job description

O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences. Our Culture Cloud is a suite of apps designed to enhance the employee experience with strategic recognition, service awards, wellbeing, leadership, and events that help people thrive at work. Our Culture by Design approach provides expert services to organizations looking to create great workplaces.

Our global team of 1,500 people hail from 58 countries and speak 62 languages. As programmers, researchers, designers, client professionals and craftspeople we create the tech, tools and awards that connect employees to purpose at thousands of companies. Join us as we help people all over the world thrive at work.

Location: Salt Lake City, UT

As the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition platform. You will build, mentor, and empower a team of Site Reliability Engineers while partnering closely with Engineering, Product, and Support organizations to deliver highly available, scalable, and resilient services that serve millions of users. We are seeking a leader who is passionate about operational excellence, continuous improvement, and fostering a reliability-first culture through automation, observability, and shared ownership. In this role, you will champion the development of self-healing platforms, drive incident and operational maturity, and enable engineering teams to innovate faster while delivering exceptional customer experiences.

Key Responsibilities:

  • Lead, mentor, and develop a team of Site Reliability Engineers, fostering a culture of reliability, accountability, operational excellence, and continuous improvement.
  • Define and execute the organization's reliability strategy, improving availability, scalability, performance, and resilience through automation and engineering best practices.
  • Establish team priorities, goals, and success metrics aligned with business objectives, customer needs, and platform health.
  • Partner with Engineering, Product, and Support leaders to drive shared ownership of production services and embed reliability, observability, and operational excellence throughout the software development lifecycle.
  • Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar tools, establishing enterprise standards for metrics, logs, traces, alerting, and Service Level Objectives (SLOs).
  • Oversee production triage, incident response, and escalation processes, ensuring timely service restoration, effective root cause analysis, and blameless post-incident reviews.
  • Champion a reliability-first engineering culture focused on automation, proactive risk reduction, operational readiness, shift-left quality practices, and continuous improvement.
  • Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective operational handoffs, and consistent service ownership.
  • Own on-call programs, incident management practices, and operational health metrics, driving improvements in alert quality, operational efficiency, and toil reduction.
  • Manage team capacity, hiring, performance management, career development, budgeting, and workforce planning to ensure effective support of business-critical services.
  • Provide regular reporting to engineering and executive leadership on reliability trends, incidents, risks, performance metrics, and strategic initiatives.

Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines, including 2+ years in a technical leadership or people management role.
  • Proven experience leading teams responsible for production operations, reliability engineering, incident management, and operational excellence.
  • Experience designing and implementing SRE practices, reliability programs, or operational maturity initiatives within growing engineering organizations.
  • Experience operating large-scale, customer-facing SaaS platforms with high availability, performance, and scalability requirements.
  • Strong understanding of modern software engineering practices and partnering with development teams to build reliable, resilient systems.
  • Hands-on experience with observability platforms such as OpenTelemetry, Datadog, Coralogix, or similar technologies.
  • Strong knowledge of AWS and Kubernetes in production environments.
  • Deep understanding of monitoring, logging, distributed tracing, SLIs, SLOs, error budgets, and reliability engineering principles.
  • Demonstrated ability to lead cross-functional initiatives and influence stakeholders across Engineering, Product, and Support organizations.
  • Experience developing engineering roadmaps, defining team objectives, aligning reliability investments with business priorities, and driving continuous operational improvement through incident learning and post-incident reviews.

Preferred Qualifications

  • Experience leading distributed or globally dispersed engineering teams.
  • Experience with multiple cloud providers or cloud-agnostic platform architectures.
  • Familiarity with security, compliance, governance, and operational risk management frameworks.
  • Proficiency with modern Infrastructure-as-Code and technologies such as Terraform, Golang, Python, Playwright, and Performance Monitoring tools.
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.
  • Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics.