Sr. Observability Engineer
Lehi, UT · On-site
Escalate to engineering managers when coverage fails or an owner is missing. * Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no ...
Lehi, UT · On-site
Escalate to engineering managers when coverage fails or an owner is missing. * Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no ...
Lehi, UT · On-site
Escalate to engineering managers when coverage fails or an owner is missing. * Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no ...
Lehi, UT · On-site
$98K - $134K/yr
Escalate to engineering managers when coverage fails or an owner is missing. * Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no ...
Lehi, UT · On-site
$98K - $134K/yr
Escalate to engineering managers when coverage fails or an owner is missing. * Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no ...
Observability Engineer (OpsRamp) - Secret clearance You will be part of a larger technical team ... Log Management: Configure log collection, aggregation, and analysis * Alerting & Notifications:
Quick apply
Observability Engineer (OpsRamp) - Secret clearance You will be part of a larger technical team ... Log Management: Configure log collection, aggregation, and analysis * Alerting & Notifications:
Observability Engineer (OpsRamp) - Secret clearance You will be part of a larger technical team ... Log Management: Configure log collection, aggregation, and analysis * Alerting & Notifications:
New
Observability Engineer (OpsRamp) - Secret clearance You will be part of a larger technical team ... Log Management: Configure log collection, aggregation, and analysis * Alerting & Notifications:
New
Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar ... Manage team capacity, hiring, performance management, career development, budgeting, and workforce ...
Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar ... Manage team capacity, hiring, performance management, career development, budgeting, and workforce ...
Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar ... Manage team capacity, hiring, performance management, career development, budgeting, and workforce ...
Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar ... Manage team capacity, hiring, performance management, career development, budgeting, and workforce ...
Salt Lake City, UT · On-site
$120 - $180/hr
Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar ... Manage team capacity, hiring, performance management, career development, budgeting, and workforce ...
Salt Lake City, UT · On-site
$120 - $180/hr
Build and evolve observability capabilities using OpenTelemetry, Datadog, Coralogix, or similar ... Manage team capacity, hiring, performance management, career development, budgeting, and workforce ...
Performance Engineering & Observability * Own the performance engineering discipline, including ... Work with Product Management to define performance requirements, benchmarks, and quality gates.
Performance Engineering & Observability * Own the performance engineering discipline, including ... Work with Product Management to define performance requirements, benchmarks, and quality gates.
Lehi, UT · On-site
... observability, and automated testing • Ensure systems are reliable, scalable, and well ... management responsibilities • Deep experience with cloud-native architectures and building ...
Lehi, UT · On-site
... observability, and automated testing • Ensure systems are reliable, scalable, and well ... management responsibilities • Deep experience with cloud-native architectures and building ...
South Jordan, UT · On-site
$100 - $130/hr
Lead and manage critical incidents, ensuring timely resolution and effective communication with ... Drive automation, toil reduction, and enhancements in observability, monitoring, and reliability ...
South Jordan, UT · On-site
$100 - $130/hr
Lead and manage critical incidents, ensuring timely resolution and effective communication with ... Drive automation, toil reduction, and enhancements in observability, monitoring, and reliability ...
Manage existing and new external vendor relationships * Oversee a team of platform engineers ... Experience with DEX (Digital Employee Experience) observability platforms such as Nexthink or ...
Manage existing and new external vendor relationships * Oversee a team of platform engineers ... Experience with DEX (Digital Employee Experience) observability platforms such as Nexthink or ...
Drive automation, toil reduction, and enhancements in observability, monitoring, and reliability ... Proven people management and team leadership experience. * Strong working knowledge of UNIX/Linux ...
Drive automation, toil reduction, and enhancements in observability, monitoring, and reliability ... Proven people management and team leadership experience. * Strong working knowledge of UNIX/Linux ...
Drive automation, toil reduction, and enhancements in observability, monitoring, and reliability ... Proven people management and team leadership experience. * Strong working knowledge of UNIX/Linux ...
Drive automation, toil reduction, and enhancements in observability, monitoring, and reliability ... Proven people management and team leadership experience. * Strong working knowledge of UNIX/Linux ...
Salt Lake City, UT · On-site
We value hands-on leadership, deep technical judgment, and the ability to manage remote and on-site collaborators. Visa sponsorship is not available; however, multiple work models are supported. #J ...
Salt Lake City, UT · On-site
We value hands-on leadership, deep technical judgment, and the ability to manage remote and on-site collaborators. Visa sponsorship is not available; however, multiple work models are supported. #J ...
Lehi, UT · On-site
Our Engineering Managers are hands-on technical leaders who bring both strategic vision and the ... Champion modern engineering practices, including cloud-native development, observability, and ...
Lehi, UT · On-site
Our Engineering Managers are hands-on technical leaders who bring both strategic vision and the ... Champion modern engineering practices, including cloud-native development, observability, and ...
Lehi, UT · On-site
Our Engineering Managers are hands-on technical leaders who bring both strategic vision and the ... Champion modern engineering practices, including cloud-native development, observability, and ...
Lehi, UT · On-site
Our Engineering Managers are hands-on technical leaders who bring both strategic vision and the ... Champion modern engineering practices, including cloud-native development, observability, and ...
Draper, UT · On-site
$112.50 - $137.50/hr
Observability:** Develop and manage robust monitoring, logging, and alerting systems to gain deep visibility into cloud infrastructure, API health, and IoT endpoint performance.* **Incident Response:
Draper, UT · On-site
$112.50 - $137.50/hr
Observability:** Develop and manage robust monitoring, logging, and alerting systems to gain deep visibility into cloud infrastructure, API health, and IoT endpoint performance.* **Incident Response:
Salt Lake City, UT · On-site +1
$130K - $160K/yr
Familiarity with common data stores and managed services (e.g., Postgres, MongoDB, DynamoDB) and how they fail in distributed systems. * Experience with at least two observability stacks (Prometheus ...
Salt Lake City, UT · On-site +1
$130K - $160K/yr
Familiarity with common data stores and managed services (e.g., Postgres, MongoDB, DynamoDB) and how they fail in distributed systems. * Experience with at least two observability stacks (Prometheus ...
Salt Lake City, UT · On-site
$130K - $160K/yr
Familiarity with common data stores and managed services (e.g., Postgres, MongoDB, DynamoDB) and how they fail in distributed systems. * Experience with at least two observability stacks (Prometheus ...
Salt Lake City, UT · On-site
$130K - $160K/yr
Familiarity with common data stores and managed services (e.g., Postgres, MongoDB, DynamoDB) and how they fail in distributed systems. * Experience with at least two observability stacks (Prometheus ...
Salt Lake City, UT · On-site
$150 - $200/hr
IP address management (IPAM) * Network provisioning and lifecycle * Cloud and data center network integration * Leverage NetBox for IPAM, DCIM, and infrastructure modeling Monitoring, Observability ...
Salt Lake City, UT · On-site
$150 - $200/hr
IP address management (IPAM) * Network provisioning and lifecycle * Cloud and data center network integration * Leverage NetBox for IPAM, DCIM, and infrastructure modeling Monitoring, Observability ...
| Aspect | Observability Manager | Site Reliability Engineer |
|---|---|---|
| Credentials | Typically requires experience in monitoring, logging, and cloud tools; certifications like AWS, Google Cloud, or Kubernetes are common | Requires strong background in systems engineering, scripting, and cloud platforms; certifications like AWS, GCP, or Linux are often preferred |
| Work Environment | Focuses on overseeing observability tools, data analysis, and team coordination in tech environments | Hands-on role involving system automation, incident response, and infrastructure reliability |
| Industry Usage | Used across tech companies to improve system visibility and performance | Common in DevOps and SRE teams to ensure system reliability and uptime |
The Observability Manager primarily oversees monitoring and logging strategies, ensuring system visibility, while the Site Reliability Engineer is more hands-on, focusing on automating infrastructure and maintaining system reliability. Both roles require technical expertise and often collaborate closely but differ in scope and daily responsibilities.
The most popular types of Observability jobs in Utah are:
For Observability Manager jobs in Utah, the most frequently searched job titles are:
The top searched job categories for Observability Manager jobs in Utah are:
Cities in Utah with the most Observability Manager job openings:
At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime.
We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand.
We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric.
This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.
Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL and Redis. Datadog is our observability platform and incident.io is our incident response platform.
What you'll do:Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.
Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.
Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.
Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.
Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.
Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.
Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.
After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.
Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.
BS in Computer Science or equivalent experience
5+ years running production observability, SRE, or DevOps
5+ years automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency
AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs
Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis
Shared on-call, Incident Commander-capable
Fintech experience with MX-like architectures
Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast
Google SRE practices: toil elimination, incident management, automation for self-healing
Cross-functional influence without authority. You've improved teams that don't report to you
Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends)
OpenTelemetry instrumentation
Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience
Golang and Ruby on Rails (the MX stack)
Sourced by ZipRecruiter