1

Observability Manager Jobs in Utah (NOW HIRING)

Senior DevOps Engineer I

American Fork, UT · On-site

$116K - $149K/yr

Experience or exposure to IoT/edge fleet management and MQTT brokers , observability platforms like Datadog or Coralogix , AWS GovCloud/FedRAMP compliance contexts , or overlay networking like ...

Sr. DevOps Engineer

Salt Lake City, UT · On-site

$125K - $161K/yr

Ability to work in a team and manage multiple tasks simultaneously. * Previous experience in a DevOps or related role is preferred. * Solid understanding of observability tools and techniques (e.g ...

API Design & Ecosystem Management: Lead the design and implementation of resilient RESTful and ... Observability & Incident Response: Implement robust observability, telemetry, and distributed ...

Senior Data Engineer

South Jordan, UT · On-site

$100K - $137K/yr

Manage the interface between our AI systems and the input data they need to operate via streaming ... metrics, and observability systems Requirements What you bring: * Good to Expert level ...

Sr. Site Reliability Engineer

Salt Lake City, UT · On-site

$55.25 - $73.25/hr

Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar tools ... Drive incident management, root cause analysis, and blameless post incident reviews to improve ...

Senior Data Engineer

Lehi, UT · On-site

$99K - $135K/yr

Data Quality & Observability Platforms * Metadata Management Solutions ABOUT WAYSTAR Through a smart platform and better experience, Waystar helps providers simplify healthcare payments and yield ...

Senior Data Engineer

Lehi, UT · On-site

$99K - $135K/yr

Data Quality & Observability Platforms * Metadata Management Solutions ABOUT WAYSTAR Through a smart platform and better experience, Waystar helps providers simplify healthcare payments and yield ...

API Design & Ecosystem Management: Lead the design and implementation of resilient RESTful and ... Observability & Incident Response: Implement robust observability, telemetry, and distributed ...

Senior Software Engineer - Backend

American Fork, UT · On-site

$109K - $144K/yr

API Design & Ecosystem Management: Lead the design and implementation of resilient RESTful and ... Observability & Incident Response: Implement robust observability, telemetry, and distributed ...

Platform Engineer (DevOps)

Salt Lake City, UT · On-site

$52 - $71.25/hr

Administer enterprise Source Code Management (SCM) repositories, implementing branching strategies ... Implement enterprise monitoring, logging, and observability solutions using Splunk, AppDynamics ...

Showing results 21-40

Observability Manager information

What is the difference between Observability Manager vs Site Reliability Engineer?

AspectObservability ManagerSite Reliability Engineer
CredentialsTypically requires experience in monitoring, logging, and cloud tools; certifications like AWS, Google Cloud, or Kubernetes are commonRequires strong background in systems engineering, scripting, and cloud platforms; certifications like AWS, GCP, or Linux are often preferred
Work EnvironmentFocuses on overseeing observability tools, data analysis, and team coordination in tech environmentsHands-on role involving system automation, incident response, and infrastructure reliability
Industry UsageUsed across tech companies to improve system visibility and performanceCommon in DevOps and SRE teams to ensure system reliability and uptime

The Observability Manager primarily oversees monitoring and logging strategies, ensuring system visibility, while the Site Reliability Engineer is more hands-on, focusing on automating infrastructure and maintaining system reliability. Both roles require technical expertise and often collaborate closely but differ in scope and daily responsibilities.

What are the most commonly searched types of Observability jobs in Utah?

The most popular types of Observability jobs in Utah are:

What are popular job titles related to Observability Manager jobs in Utah?

For Observability Manager jobs in Utah, the most frequently searched job titles are:

What cities in Utah are hiring for Observability Manager jobs?

Cities in Utah with the most Observability Manager job openings:

Senior Site Reliability Engineer (SRE)

CenCore LLC

Springville, UT • On-site, Remote

$52.75 - $70/hr

Full-time

Posted 27 days ago


Job description

Description
The Senior Site Reliability Engineer (SRE) will implement, secure, and operate the cloud infrastructure that supports CenCore Group's proprietary enterprise SaaS platform. This role is responsible for maintaining a scalable, highly available, secure, and reliable cloud environment as the platform grows and supports enterprise customers.
Key Responsibilities
  • Manage, maintain, and improve AWS-based cloud infrastructure supporting enterprise SaaS operations.
  • Operate and support Kubernetes environments, including Amazon EKS.
  • Own platform reliability, scalability, availability, disaster recovery readiness, and operational resilience.
  • Design and support cloud networking, load balancing, routing, traffic management, and related infrastructure components.
  • Implement and maintain monitoring, alerting, logging, and observability solutions to support proactive issue detection and response.
  • Establish and document operational standards, Service Level Objectives (SLOs), incident response processes, and reliability best practices.
  • Partner with software engineering and product teams to improve application performance, platform stability, and deployment reliability.
  • Apply security best practices across IAM, secrets management, encryption, vulnerability remediation, access controls, and production operations.
  • Support production operations, troubleshoot critical issues, and participate in incident resolution as needed.

Requirements
Required Qualifications
  • Active Top Secret clearance with SCI eligibility
  • Professional experience supporting cloud infrastructure, site reliability, DevOps, platform engineering, or systems engineering functions.
  • Hands-on experience with AWS cloud services and production cloud operations.
  • Experience administering or operating Kubernetes environments.
  • Working knowledge of infrastructure reliability, availability, scalability, incident response, and operational support practices.
  • Experience implementing monitoring, logging, alerting, or observability tools.
  • Ability to troubleshoot complex production issues and coordinate resolution across technical teams.
  • Strong understanding of cloud security fundamentals, including identity and access management, encryption, secrets management, and vulnerability remediation.
  • Ability to document technical processes, standards, and operational procedures.

Preferred Qualifications
  • Experience with AWS services such as EKS, ALB, VPC, CloudFront, Route 53, RDS/Aurora, S3, and IAM.
  • Experience with Terraform or other Infrastructure as Code tools.
  • Experience with monitoring platforms such as Datadog, CloudWatch, Grafana, Prometheus, or similar tools.
  • PostgreSQL administration, performance tuning, or database operations experience.
  • Experience supporting enterprise SaaS, cloud-native applications, or customer-facing production platforms.
  • Experience developing disaster recovery, operational readiness, or production support documentation.

Skills / Competencies
  • Cloud infrastructure operations and automation
  • Platform reliability, scalability, and performance optimization
  • Kubernetes administration and containerized application support
  • Monitoring, observability, and incident response
  • Cloud security and operational risk awareness
  • Technical troubleshooting and root cause analysis
  • Cross-functional collaboration with engineering, product, and operations teams
  • Clear technical documentation and process improvement

Work Environment and Physical Requirements
This role is primarily performed in a professional office or remote technology environment, depending on business needs and position requirements. Work involves regular use of a computer, collaboration tools, and cloud-based systems. The position may require participation in production support, incident response, or after-hours troubleshooting as needed. Physical requirements are generally sedentary and include prolonged periods of sitting, computer use, and communicating with internal teams.