1

Director Site Reliability Engineering Jobs in Utah

$57.75 - $76.75/hr

Site Reliability Engineer (SRE) Department: Technology Location: Manila Reporting To: Head of Infra ... You will collaborate with engineering, DevOps, and client success teams to operationalize ...

Site Reliability Engineer

Saint George, UT · On-site

$50.75 - $67.50/hr

Shares best practices - Shares understanding of Site Reliability Engineering culture across organization; shares knowledge of best practices, approaches, documentation, and code with team members and ...

Site Reliability Engineer

Saint George, UT · On-site

$53.75 - $71.50/hr

Shares best practices - Shares understanding of Site Reliability Engineering culture across organization; shares knowledge of best practices, approaches, documentation, and code with team members and ...

Sr. Site Reliability Engineer

Salt Lake City, UT · On-site

$55.25 - $73.25/hr

Salt Lake City, UT As a Senior Site Reliability Engineer, you will help define the future of ... You'll leverage software engineering, automation, and cloud-native technologies to build and ...

Salt Lake City, UT As a Senior Site Reliability Engineer, you will help define the future of ... You'll leverage software engineering, automation, and cloud-native technologies to build and ...

Site Reliability Engineer III - Neovest Orem, UT, United States Job Information * Job ... Job Category Software Engineering * Business Unit Commercial & Investment Bank * Posting Date 09/06 ...

New

Staff SRE

Pleasant Grove, UT · On-site

$51.50 - $68.25/hr

The Staff Site Reliability Engineer will architect and implement enterprise-scale infrastructure ... leadership across engineering teams to ensure exceptional system reliability and performance.

Site Reliability Engineer

Draper, UT · On-site

$53.25 - $70.75/hr

... hardware engineering teams to define Service Level Indicators (SLIs), Service Level Objectives ... an SRE, DevOps, or Infrastructure Engineering role supporting production cloud environments • ...

Senior Site Reliability Engineer

Lehi, UT

$53.50 - $71/hr

This role bridges software engineering and cloud operations, ensuring mission-critical systems remain highly available and resilient. By integrating reliability early, the SRE fosters a culture of ...

Site Reliability Engineer III - Neovest

Orem, UT · On-site

$49.50 - $65.75/hr

Supports the adoption of site reliability engineering best practices within your team Required qualifications, capabilities, and skills * Formal training or certification on software engineering ...

Senior Site Reliability Engineer

Lehi, UT · On-site

$53.50 - $71/hr

This role bridges software engineering and cloud operations, ensuring mission-critical systems remain highly available and resilient. By integrating reliability early, the SRE fosters a culture of ...

Staff SRE

Pleasant Grove, UT · On-site

$51.50 - $68.25/hr

Staff Site Reliability Engineer Join Us at Pura--Reimagining Fragrance for the Future At Pura, we ... Partner with engineering leadership to drive reliability improvements through advanced automated ...

next page

Showing results 1-20

Director Site Reliability Engineering information

See Utah salary details

$9

$58

$83

How much do director site reliability engineering jobs pay per hour?

As of Aug 8, 2026, the average hourly pay for director site reliability engineering in Utah is $58.03, according to ZipRecruiter salary data. Most workers in this role earn between $49.90 and $66.30 per hour, depending on experience, location, and employer.

How much do director site reliability engineering get paid?

Director of Site Reliability Engineering typically earns a salary ranging from $150,000 to $250,000 annually, depending on experience, company size, and location. They often oversee teams using tools like Kubernetes and Prometheus and require strong leadership and technical skills.

What is a director site reliability engineering?

A Director of Site Reliability Engineering (SRE) leads teams responsible for ensuring the availability, performance, and scalability of software systems. They define reliability best practices, drive automation, and collaborate with engineering and product teams to improve system resilience. This role requires strong leadership, technical expertise, and a focus on balancing innovation with operational stability.

What are the main challenges faced by a director site reliability engineering, and how can I prepare for them?

A Director of Site Reliability Engineering often encounters challenges such as balancing rapid feature delivery with system stability, managing complex incident responses, and fostering a culture of continuous improvement. Additionally, aligning reliability goals with business objectives and securing cross-functional buy-in can be demanding. To prepare, it is helpful to gain experience in high-scale system management, develop strong leadership and communication abilities, and cultivate a proactive approach to risk management and automation. Staying up to date with the latest SRE practices and building relationships with both engineering and business teams will also support your success in this pivotal role.

What are the key skills and qualifications needed to thrive as a director site reliability engineering?

To thrive as a Director Site Reliability Engineering, you need extensive experience in software engineering, infrastructure management, incident response, and people leadership, often supported by a degree in computer science or a related field. Familiarity with cloud platforms (such as AWS, GCP, or Azure), automation tools (Terraform, Ansible), monitoring systems (Prometheus, Datadog), and relevant certifications like CKA or AWS Solutions Architect is valued. Outstanding communication, stakeholder management, and strategic vision are key soft skills that set leaders apart in this role. These abilities ensure the reliability, scalability, and efficiency of critical systems while effectively guiding and motivating technical teams.

What cities in Utah are hiring for Director Site Reliability Engineering jobs? Cities in Utah with the most Director Site Reliability Engineering job openings:
Infographic showing various Director Site Reliability Engineering job openings in Utah as of July 2026, with employment types broken down into 86% Full Time, and 14% Part Time. Highlights an 76% In-person, 4% Hybrid, and 20% Remote job distribution, with an average salary of $120,700 per year, or $58 per hour.

Manager, Site Reliability Engineering

OC Tanner

Salt Lake City, UT • On-site

Full-time

Posted 18 days ago


Job description

O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences. Our Culture Cloud is a suite of apps designed to enhance the employee experience with strategic recognition, service awards, wellbeing, leadership, and events that help people thrive at work. Our Culture by Design approach provides expert services to organizations looking to create great workplaces.
Our global team of 1,500 people hail from 58 countries and speak 62 languages. As programmers, researchers, designers, client professionals and craftspeople we create the tech, tools and awards that connect employees to purpose at thousands of companies. Join us as we help people all over the world thrive at work.
Location: Salt Lake City, UT
As the Manager of Site Reliability Engineering, you will lead the strategy, execution, and evolution of reliability for our world-class employee recognition platform. You will build, mentor, and empower a team of Site Reliability Engineers while partnering closely with Engineering, Product, and Support organizations to deliver highly available, scalable, and resilient services that serve millions of users. We are seeking a leader who is passionate about operational excellence, continuous improvement, and fostering a reliability-first culture through automation, observability, and shared ownership. In this role, you will champion the development of self-healing platforms, drive incident and operational maturity, and enable engineering teams to innovate faster while delivering exceptional customer experiences.
Key Responsibilities:
  • Lead, mentor, and develop a team of Site Reliability Engineers, fostering a culture of reliability, accountability, operational excellence, and continuous improvement.
  • Define and execute the organization's reliability strategy, improving availability, scalability, performance, and resilience through automation and engineering best practices.
  • Establish team priorities, goals, and success metrics aligned with business objectives, customer needs, and platform health.
  • Partner with Engineering, Product, and Support leaders to drive shared ownership of production services and embed reliability, observability, and operational excellence throughout the software development lifecycle.
  • Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar tools, establishing enterprise standards for metrics, logs, traces, alerting, and Service Level Objectives (SLOs).
  • Oversee production triage, incident response, and escalation processes, ensuring timely service restoration, effective root cause analysis, and blameless post-incident reviews.
  • Champion a reliability-first engineering culture focused on automation, proactive risk reduction, operational readiness, shift-left quality practices, and continuous improvement.
  • Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective operational handoffs, and consistent service ownership.
  • Own on-call programs, incident management practices, and operational health metrics, driving improvements in alert quality, operational efficiency, and toil reduction.
  • Manage team capacity, hiring, performance management, career development, budgeting, and workforce planning to ensure effective support of business-critical services.
  • Provide regular reporting to engineering and executive leadership on reliability trends, incidents, risks, performance metrics, and strategic initiatives.

Required Qualifications
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines, including 2+ years in a technical leadership or people management role.
  • Proven experience leading teams responsible for production operations, reliability engineering, incident management, and operational excellence.
  • Experience designing and implementing SRE practices, reliability programs, or operational maturity initiatives within growing engineering organizations.
  • Experience operating large-scale, customer-facing SaaS platforms with high availability, performance, and scalability requirements.
  • Strong understanding of modern software engineering practices and partnering with development teams to build reliable, resilient systems.
  • Hands-on experience with observability platforms such as OpenTelemetry, Datadog, Coralogix, or similar technologies.
  • Strong knowledge of AWS and Kubernetes in production environments.
  • Deep understanding of monitoring, logging, distributed tracing, SLIs, SLOs, error budgets, and reliability engineering principles.
  • Demonstrated ability to lead cross-functional initiatives and influence stakeholders across Engineering, Product, and Support organizations.
  • Experience developing engineering roadmaps, defining team objectives, aligning reliability investments with business priorities, and driving continuous operational improvement through incident learning and post-incident reviews.

Preferred Qualifications
  • Experience leading distributed or globally dispersed engineering teams.
  • Experience with multiple cloud providers or cloud-agnostic platform architectures.
  • Familiarity with security, compliance, governance, and operational risk management frameworks.
  • Proficiency with modern Infrastructure-as-Code and technologies such as Terraform, Golang, Python, Playwright, and Performance Monitoring tools.
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.
  • Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics.