1

Observability Site Reliability Engineer Jobs in Utah

Sr. Site Reliability Engineer

Lehi, UT ยท On-site

$53.50 - $71/hr

As a Senior Observability Engineer, you build and operate an observability control plane. You ... Google SRE practices: toil elimination, incident management, automation for self-healing * Cross ...

New

Sr. Site Reliability Engineer

Lehi, UT ยท On-site

$53.50 - $71/hr

As a Senior Observability Engineer, you build and operate an observability control plane. You ... Google SRE practices: toil elimination, incident management, automation for self-healing * Cross ...

New

Sr. Site Reliability Engineer

Salt Lake City, UT ยท On-site

$55.25 - $73.25/hr

Salt Lake City, UT As a Senior Site Reliability Engineer, you will help define the future of ... Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar tools.

Salt Lake City, UT As a Senior Site Reliability Engineer, you will help define the future of ... Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar tools.

Senior Site Reliability Engineer

Lehi, UT

$53.50 - $71/hr

Job summary The Site Reliability Engineer (SRE) collaborates with development teams to embed ... Implement observability and automation-first principles to measure system health and drive ...

Staff SRE

Pleasant Grove, UT ยท On-site

$51.50 - $68.25/hr

The Staff Site Reliability Engineer will architect and implement enterprise-scale infrastructure ... observability platforms and practices, including implementing custom monitoring solutions and ...

$57.75 - $76.75/hr

Site Reliability Engineer (SRE) Department: Technology Location: Manila Reporting To: Head of Infra Tookitaki is looking for a Site Reliability Engineer (SRE) with 3-6 years of experience to help ...

Senior Site Reliability Engineer

Lehi, UT ยท On-site

$53.50 - $71/hr

By integrating reliability early, the SRE fosters a culture of shared responsibility while enabling ... Implement observability and automation-first principles to measure system health and drive ...

Site Reliability Engineer

Saint George, UT ยท On-site

$50.75 - $67.50/hr

The Site Reliability Engineer works as part of a team to analyze, troubleshoot, deploy, monitor, and maintain TCN's large production environment with global scale. These significant responsibilities ...

Site Reliability Engineer

Saint George, UT ยท On-site

$53.75 - $71.50/hr

The Site Reliability Engineer works as part of a team to analyze, troubleshoot, deploy, monitor, and maintain TCN's large production environment with global scale. These significant responsibilities ...

Staff SRE

Pleasant Grove, UT ยท On-site

$51.50 - $68.25/hr

Staff Site Reliability Engineer Join Us at Pura--Reimagining Fragrance for the Future At Pura, we ... Extensive experience with advanced observability platforms and practices, including implementing ...

Staff SRE

Pleasant Grove, UT ยท On-site

$51.50 - $68.25/hr

Staff Site Reliability Engineer Join Us at Pura-Reimagining Fragrance for the Future At Pura, we ... Extensive experience with advanced observability platforms and practices, including implementing ...

Site Reliability Engineer

Draper, UT ยท On-site

$53.25 - $70.75/hr

They are seeking a hands-on Site Reliability Engineer to bridge the gap between software development and infrastructure operations, ensuring the reliability and performance of their platform.

next page

Showing results 1-20

Observability Site Reliability Engineer information

What engineer makes $500,000 a year?

A senior or principal Site Reliability Engineer (SRE) or Observability Engineer with extensive experience, specialized skills, and working at large tech companies can earn $500,000 or more annually. Compensation often includes base salary, bonuses, and stock options, especially in high-demand markets and organizations with complex infrastructure.

Is AI replacing SRE?

AI is augmenting the work of Site Reliability Engineers (SREs) by automating tasks such as monitoring, incident detection, and response. However, SREs are still essential for designing systems, managing complex issues, and making strategic decisions that require human judgment. AI tools are considered complementary rather than replacements for SREs' expertise and problem-solving skills.

What engineers make $200,000 a year?

Senior Site Reliability Engineers and Observability Engineers with extensive experience, advanced skills in cloud platforms, automation, and monitoring tools can earn $200,000 or more annually. High compensation often correlates with working at large tech companies, possessing specialized certifications, and managing complex, scalable systems.

What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?

AspectObservability Site Reliability EngineerMonitoring Engineer
FocusEnsuring system reliability through observability, automation, and incident responseImplementing and managing monitoring tools and dashboards
SkillsCloud platforms, scripting, incident management, observability toolsMonitoring tools, alerting systems, data analysis
Work EnvironmentDevOps teams, cloud infrastructure, large-scale systemsOperations teams, infrastructure monitoring

While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.

What engineers make $300,000 a year?

Senior Site Reliability Engineers and Observability Engineers with extensive experience, advanced skills in cloud platforms, automation, and monitoring tools can earn $300,000 or more annually. High compensation often correlates with working at large tech companies, possessing specialized certifications, and taking on leadership or highly technical roles.
What job categories do people searching Observability Site Reliability Engineer jobs in Utah look for? The top searched job categories for Observability Site Reliability Engineer jobs in Utah are:
What cities in Utah are hiring for Observability Site Reliability Engineer jobs? Cities in Utah with the most Observability Site Reliability Engineer job openings:

Sr. Site Reliability Engineer

MX Technologies, Inc.

Lehi, UT โ€ข On-site

$53.50 - $71/hr

Other

Posted yesterday

New


Job description

At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime.

We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand.

We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric.

This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.

Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL and Redis. Datadog is our observability platform and incident.io is our incident response platform.

What you'll do:
  • Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.

  • After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.

  • Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.

  • Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.

  • Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.

  • Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.

  • Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.

  • Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.

  • Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.

  • After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.

  • Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.

Basic Requirements
  • BS in Computer Science or equivalent experience

  • 5+ years running production observability, SRE, or DevOps

  • 5+ years automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency

  • AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs

  • Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis

  • Shared on-call, Incident Commander-capable

Preferred Requirements
  • Fintech experience with MX-like architectures

  • Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast

  • Google SRE practices: toil elimination, incident management, automation for self-healing

  • Cross-functional influence without authority. You've improved teams that don't report to you

  • Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends)

  • OpenTelemetry instrumentation

  • Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience

  • Golang and Ruby on Rails (the MX stack)