1

Linux Site Reliability Engineer Jobs in Orem, UT

Sr. Observability Engineer

Lehi, UT · On-site

$98K - $134K/yr

Google SRE practices: toil elimination, incident management, automation for self-healing * Cross-functional influence without authority. You've improved teams that don't report to you * Governance ...

AI DevOps Engineer

Sandy, UT · On-site

$50.25 - $68.75/hr

... SRE, and Infrastructure teams to deliver end-to-end capabilities Improve system performance, reliability, and observability Key Responsibilities Design and develop scalable backend systems for AI ...

Partner across engineering, product, SRE, security, and finance to align priorities, influence roadmaps, and resolve cross-organizational dependencies that impact delivery at scale. * Establish and ...

AI DevOps Engineer

Sandy, UT · On-site

$50.25 - $68.75/hr

... SRE, and Infrastructure teams to deliver end-to-end capabilities Improve system performance, reliability, and observability Key Responsibilities Design and develop scalable backend systems for AI ...

Partner with SRE/Platform teams to improve CI/CD, infrastructure-as-code, and reliability practices (e.g., automated rollouts, canarying, and repeatable environments) What You Will Need to Accomplish ...

Experience: 5+ years (for Senior level) of professional experience in roles such as Performance Engineer, Software Engineer in Test, SRE, DevOps, or Software Engineer with a deep focus on the ...

AI DevOps Engineer

Sandy, UT · On-site

$50.25 - $68.75/hr

... SRE, and Infrastructure teams to deliver end-to-end capabilities • Improve system performance, reliability, and observability Key Responsibilities • Design and develop scalable backend systems ...

Sr. Platform Engineer - Data Infrastructure

Lehi, UT · On-site

$101K - $138K/yr

A background in operations such as networking, system administration, or DevOps/SRE * Enjoys working in a greenfield environment using rapid prototyping Employment with Weave is contingent upon the ...

Showing results 41-60

Linux Site Reliability Engineer information

See Orem, UT salary details

$9

$55

$79

How much do linux site reliability engineer jobs pay per hour?

As of Aug 19, 2026, the average hourly pay for linux site reliability engineer in Orem, UT is $55.42, according to ZipRecruiter salary data. Most workers in this role earn between $47.64 and $63.32 per hour, depending on experience, location, and employer.

What is a Linux Site Reliability Engineer?

A Linux Site Reliability Engineer (SRE) is an IT professional responsible for ensuring the reliability, scalability, and performance of systems running on the Linux operating system. They bridge the gap between software development and operations by automating processes, monitoring infrastructure, and managing incidents. Linux SREs focus on system availability, building tools for deployment and monitoring, and improving system robustness through best practices and automation. Their work helps organizations deliver reliable online services and quickly recover from outages or system failures.

What are the key skills and qualifications needed to thrive as a Linux Site Reliability Engineer?

To thrive as a Linux Site Reliability Engineer, you need deep expertise in Linux system administration, scripting (such as Bash or Python), and a solid understanding of networking concepts, usually backed by a computer science degree or equivalent experience. Familiarity with configuration management tools (like Ansible, Puppet, or Chef), containerization (Docker, Kubernetes), and cloud platforms (AWS, GCP, or Azure) is typically required, along with relevant certifications like RHCE or AWS Certified SysOps Administrator. Strong problem-solving skills, effective communication, and the ability to work under pressure are crucial soft skills for this role. These competencies ensure the reliability, scalability, and security of complex infrastructure, minimizing downtime and supporting seamless operations.

What are some common challenges faced by Linux Site Reliability Engineers when scaling infrastructure, and how can they be addressed?

Linux Site Reliability Engineers often encounter challenges related to maintaining system stability and performance as infrastructure scales. Issues such as configuration drift, automation bottlenecks, and monitoring gaps can arise when managing numerous servers or services. Addressing these challenges typically involves implementing robust configuration management tools, investing in automated deployment pipelines, and enhancing observability through comprehensive monitoring and alerting solutions. Collaboration with development and operations teams is essential to ensure that scalability solutions align with business needs and technical requirements.

What is the difference between Linux Site Reliability Engineer vs Linux DevOps Engineer?

AspectLinux Site Reliability EngineerLinux DevOps Engineer
CredentialsLinux certifications, SRE-specific trainingLinux certifications, DevOps tools certifications
Work EnvironmentFocus on system reliability, monitoring, incident responseFocus on automation, CI/CD pipelines, deployment
Employer & IndustryTech companies, cloud providers, large enterprisesStartups, tech firms, software development teams
Search & Comparison IntentUnderstanding reliability roles, incident managementAutomation, deployment, continuous integration

While both roles involve Linux expertise, a Linux Site Reliability Engineer primarily focuses on maintaining system reliability, monitoring, and incident response. In contrast, a Linux DevOps Engineer emphasizes automation, continuous integration, and deployment processes. Both roles require Linux skills and often overlap, but their core responsibilities differ based on organizational needs.

What are popular job titles related to Linux Site Reliability Engineer jobs in Orem, UT?

For Linux Site Reliability Engineer jobs in Orem, UT, the most frequently searched job titles are:

What job categories do people searching Linux Site Reliability Engineer jobs in Orem, UT look for?

The top searched job categories for Linux Site Reliability Engineer jobs in Orem, UT are:

Sr. Observability Engineer

MX Technologies, Inc.

Lehi, UT • On-site

$98K - $134K/yr

Full-time

Posted 17 days ago


Job description

MX is a fintech company on a mission to empower the world to be financially strong. We build technology that helps banks, credit unions, and fintechs deliver smarter, more intuitive financial experiences to millions of people.
Like many startups, we've navigated real growth challenges - and we've come out stronger on the other side. Today, MX is in a phase ofrenewed momentum and scale, with a solid foundation and a clear vision for what's next. This is a place where thoughtful execution matters, innovation is encouraged, and individuals have real ownership over their work.
Our culture values curiosity, accountability, and impact. We give people the space to question assumptions, design better solutions, and help shape how the company grows. If you're looking to do meaningful work, influence outcomes, and grow alongside a company that's ready to move fast, you'll feel at home at MX.
At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime.
We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand.
We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric.
This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.
Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL and Redis. Datadog is our observability platform and incident.io is our incident response platform.
What you'll do:
  • Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
  • After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
  • Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.
  • Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.
  • Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.
  • Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.
  • Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.
  • Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.
  • Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.
  • After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.
  • Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.

Basic Requirements
  • BS in Computer Science or equivalent experience
  • 5+ years running production observability, SRE, or DevOps
  • 5+ years automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency
  • AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs
  • Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis
  • Shared on-call, Incident Commander-capable

Preferred Requirements
  • Fintech experience with MX-like architectures
  • Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast
  • Google SRE practices: toil elimination, incident management, automation for self-healing
  • Cross-functional influence without authority. You've improved teams that don't report to you
  • Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends)
  • OpenTelemetry instrumentation
  • Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience
  • Golang and Ruby on Rails (the MX stack)

At MX, we are a high-performance organization that thrives on trust and results. This role is based in Lehi, Utah. We believe in empowering our team members to deliver exceptional outcomes while taking advantage of our incredible office space when it best supports their work. Our Utah office features onsite perks such as company-paid meals, a sports simulator, gym, mother's lounge, and meditation room and meaningful interactions with amazing people. We encourage team members to come together in the office to collaborate, kick off key projects, or strategize cross-functionally, fostering connection and innovation.
MX is proudly committed to recruiting and retaining a diverse and inclusive workforce. As an Equal Opportunity Employer, we never discriminate based on race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, military or veteran status, status as an individual with a disability, or other applicable legally protected characteristics. We particularly welcome applications from veterans and military spouses. All your information will be kept confidential according to EEO guidelines. You may request reasonable accommodations by sending an email to hr@mx.com.