Google SRE practices: toil elimination, incident management, automation for self-healing * Cross-functional influence without authority. You've improved teams that don't report to you * Governance ...
Google SRE practices: toil elimination, incident management, automation for self-healing * Cross-functional influence without authority. You've improved teams that don't report to you * Governance ...
Test and Reliability Engineer IV
UT · On-site
$70 - $80/hr
Test and Reliability Engineer IV Location: 515 Colorow Drive, Salt Lake City, UT 84108 Contract Duration: 12 months Pay Rate: $70-80/hr (all-inclusive) Position Summary The Test and Reliability ...
Test and Reliability Engineer IV
UT · On-site
$70 - $80/hr
Test and Reliability Engineer IV Location: 515 Colorow Drive, Salt Lake City, UT 84108 Contract Duration: 12 months Pay Rate: $70-80/hr (all-inclusive) Position Summary The Test and Reliability ...
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Quick apply
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Platform Engineer/ DevOps
Midvale, UT · On-site
$49.75 - $68.25/hr
... Site Reliability Engineer (SRE), Cloud Engineer, Pipeline Engineer, or in a similar enterprise engineering role. * Strong programming and automation experience using one or more of the following:
New
Quick apply
Platform Engineer/ DevOps
Midvale, UT · On-site
$49.75 - $68.25/hr
... Site Reliability Engineer (SRE), Cloud Engineer, Pipeline Engineer, or in a similar enterprise engineering role. * Strong programming and automation experience using one or more of the following:
New
Platform Engineer (DevOps)
Salt Lake City, UT · On-site
$51 - $70/hr
... Site Reliability Engineer (SRE), Cloud Engineer, Pipeline Engineer, or in a similar enterprise engineering role. * Strong programming and automation experience using one or more of the following:
Platform Engineer (DevOps)
Salt Lake City, UT · On-site
$51 - $70/hr
... Site Reliability Engineer (SRE), Cloud Engineer, Pipeline Engineer, or in a similar enterprise engineering role. * Strong programming and automation experience using one or more of the following:
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
IT Service Management Administrator
Salt Lake City, UT · On-site
$39.95 - $66.59/hr
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
IT Service Management Administrator
Salt Lake City, UT · On-site
$39.95 - $66.59/hr
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
IT Service Management Administrator
$83K - $138K/yr
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
IT Service Management Administrator
$83K - $138K/yr
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
IT Service Management Administrator
$83K - $138K/yr
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
IT Service Management Administrator
$83K - $138K/yr
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
IT Service Management Administrator
Salt Lake City, UT · On-site
$39.95 - $66.59/hr
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
IT Service Management Administrator
Salt Lake City, UT · On-site
$39.95 - $66.59/hr
Partner with Service Desk, NOC, SRE, Infrastructure, Security, and Technology leadership to ensure operational processes remain effective, measurable, and aligned with business objectives.
Platform Engineer/ DevOps
Midvale, UT · On-site
$49.75 - $68.25/hr
Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or a related technical discipline 3+ years of experience as a Platform Engineer, DevOps Engineer, Site Reliability ...
New
Platform Engineer/ DevOps
Midvale, UT · On-site
$49.75 - $68.25/hr
Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or a related technical discipline 3+ years of experience as a Platform Engineer, DevOps Engineer, Site Reliability ...
New
Maintenance Reliability Engineer 2/3
$100K - $125K/yr
Candidate must be willing to travel to all 4 campuses and be on site. Description & responsibilities: As a Maintenance Reliability Engineer, you will specialize in various types of automated ...
Maintenance Reliability Engineer 2/3
$100K - $125K/yr
Candidate must be willing to travel to all 4 campuses and be on site. Description & responsibilities: As a Maintenance Reliability Engineer, you will specialize in various types of automated ...
Maintenance Reliability Engineer 2/3
Magna, UT · On-site
$100K - $125K/yr
Candidate must be willing to travel to all 4 campuses and be on site. Description & responsibilities: As a Maintenance Reliability Engineer, you will specialize in various types of automated ...
Maintenance Reliability Engineer 2/3
Magna, UT · On-site
$100K - $125K/yr
Candidate must be willing to travel to all 4 campuses and be on site. Description & responsibilities: As a Maintenance Reliability Engineer, you will specialize in various types of automated ...
Manager for DevOps and DBA Engineer - Assistant Vice President
Salt Lake City, UT · Remote
$150K - $180K/yr
Bachelor's degree in computer science information systems engineering or related field or equivalent experience * 8+ years DevOps cloud SRE or infrastructure engineering supporting PaaS and SaaS ...
Manager for DevOps and DBA Engineer - Assistant Vice President
Salt Lake City, UT · Remote
$150K - $180K/yr
Bachelor's degree in computer science information systems engineering or related field or equivalent experience * 8+ years DevOps cloud SRE or infrastructure engineering supporting PaaS and SaaS ...
Manager for DevOps and DBA Engineer - Assistant Vice President
Salt Lake City, UT · On-site
$150K - $180K/yr
Bachelor's degree in computer science information systems engineering or related field or equivalent experience * 8+ years DevOps cloud SRE or infrastructure engineering supporting PaaS and SaaS ...
Manager for DevOps and DBA Engineer - Assistant Vice President
Salt Lake City, UT · On-site
$150K - $180K/yr
Bachelor's degree in computer science information systems engineering or related field or equivalent experience * 8+ years DevOps cloud SRE or infrastructure engineering supporting PaaS and SaaS ...
Senior Cloud Engineer
Salt Lake City, UT · On-site
$54 - $72/hr
Serve as the central point of contact, coordinating across software engineering, SRE, and infrastructure teams to remove blockers and solve complex architectural problems. • Internal Advocacy:
Senior Cloud Engineer
Salt Lake City, UT · On-site
$54 - $72/hr
Serve as the central point of contact, coordinating across software engineering, SRE, and infrastructure teams to remove blockers and solve complex architectural problems. • Internal Advocacy:
DevOps/ Platform Engineer
Midvale, UT · On-site
$49.75 - $68.25/hr
... Site Reliability Engineering (SRE), Infrastructure Engineering, Cloud Engineering, or a related role Strong programming and scripting experience using one or more of the following: Python Java ...
DevOps/ Platform Engineer
Midvale, UT · On-site
$49.75 - $68.25/hr
... Site Reliability Engineering (SRE), Infrastructure Engineering, Cloud Engineering, or a related role Strong programming and scripting experience using one or more of the following: Python Java ...
Site Reliability Engineer information
See Draper, UT salary details
$10.11 - $17
1% of jobs
$17 - $23.88
0% of jobs
$23.88 - $30.77
0% of jobs
$30.77 - $37.65
2% of jobs
$37.65 - $44.54
4% of jobs
$44.54 - $51.42
17% of jobs
$51.49 is the 25th percentile. Wages below this are outliers.
$51.42 - $58.31
27% of jobs
$58.31 - $65.19
20% of jobs
$66.37 is the 75th percentile. Wages above this are outliers.
$65.19 - $72.07
17% of jobs
$72.07 - $78.96
6% of jobs
$78.96 - $85.84
4% of jobs
$10
$59
$85
How much do site reliability engineer jobs pay per hour?
Is a site reliability engineer a stressful job?
What is a site reliability engineer?
A site reliability engineer specializes in site reliability engineering, or SRE, a specific branch of operations first pioneered by Google. You are responsible for ensuring that when a website decides to scale a particular feature for various users to access, it does not break the underlying software or website functions. This means you need to use analytical problem-solving skills to determine how to make specific features on a new software release work on top of existing source code.
What are the key skills and qualifications needed to thrive as a site reliability engineer?
What are some of the most common challenges site reliability engineers face when balancing system reliability with rapid software delivery?
What is the difference between Site Reliability Engineer vs DevOps Engineer?
| Aspect | Site Reliability Engineer | DevOps Engineer |
|---|---|---|
| Credentials | Typically requires a computer science degree, certifications like AWS, Google Cloud, or Kubernetes | Similar credentials, often with cloud certifications and scripting skills |
| Work Environment | Focuses on maintaining and improving system reliability, often in large-scale production environments | Works on automation, CI/CD pipelines, and deployment processes across development and operations teams |
| Industry Usage | Common in tech, cloud services, and large-scale enterprise companies | Widely used in software development, cloud, and IT organizations |
Both roles require strong technical skills and cloud knowledge, but SREs focus more on system reliability and uptime, while DevOps engineers emphasize automation and deployment processes. They often collaborate but have distinct primary responsibilities.
What is a site reliability engineer?

Job description
At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime.
We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand.
We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric.
This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.
Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL and Redis. Datadog is our observability platform and incident.io is our incident response platform.
What you'll do:Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.
Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.
Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.
Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.
Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.
Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.
Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.
After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.
Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.
BS in Computer Science or equivalent experience
5+ years running production observability, SRE, or DevOps
5+ years automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency
AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs
Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis
Shared on-call, Incident Commander-capable
Fintech experience with MX-like architectures
Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast
Google SRE practices: toil elimination, incident management, automation for self-healing
Cross-functional influence without authority. You've improved teams that don't report to you
Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends)
OpenTelemetry instrumentation
Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience
Golang and Ruby on Rails (the MX stack)
About MX Technologies
Sourced by ZipRecruiter