1

Site Reliability Engineer Jobs in Sandy, UT (NOW HIRING)

Reliability Engineer

Salt Lake City, UT

$98K - $123K/yr

Job Title Reliability Engineer Summary Reliability, Maintenance, and Engineering (RME) is hiring ... Lead the Root Cause Analysis and Permanent Corrective Action activities at the site. * Identify ...

Principal Reliability Engineer

Provo, UT

$97K - $122K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

US-AZ-TUCSON-801 ~ 1151 E Hermans Rd ~ BLDG 801 (External Site) Position Role Type: Onsite U.S ... Coach other Reliability engineers and work with functional leadership in identifying and ...

Reliability Engineer

Salt Lake City, UT ยท On-site

$99K - $125K/yr

Requires collaboration with cross-functional teams, including engineering, maintenance, operations, and quality, to achieve optimal equipment reliability and production efficiency. Requirements ...

Senior Principal Reliability Engineer

Provo, UT

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

US-AZ-TUCSON-801 ~ 1151 E Hermans Rd ~ BLDG 801 (External Site) Position Role Type: Onsite U.S ... The Senior Principal Reliability Engineer will develop and implement Reliability engineering ...

Senior Reliability Engineer

Provo, UT

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

US-AZ-TUCSON-801 ~ 1151 E Hermans Rd ~ BLDG 801 (External Site) Position Role Type: Onsite U.S ... We are looking for a Senior Reliability Engineer to join our team located in Tucson, AZ. What You ...

Reliability Engineer II

Provo, UT

$97K - $122K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

US-AZ-TUCSON-801 ~ 1151 E Hermans Rd ~ BLDG 801 (External Site) Position Role Type: Onsite U.S ... We are looking for a Reliability Engineer II to join our team located in Tucson, AZ. What You Will ...

Test and Reliability Engineer IV

UT ยท On-site

$70 - $80/hr

Test and Reliability Engineer IV Location: 515 Colorow Drive, Salt Lake City, UT 84108 Contract Duration: 12 months Pay Rate: $70-80/hr (all-inclusive) Position Summary The Test and Reliability ...

Showing results 21-40

Site Reliability Engineer information

See Sandy, UT salary details

$10

$60

$87

How much do site reliability engineer jobs pay per hour?

As of Aug 14, 2026, the average hourly pay for site reliability engineer in Sandy, UT is $60.57, according to ZipRecruiter salary data. Most workers in this role earn between $52.07 and $69.23 per hour, depending on experience, location, and employer.

Is a site reliability engineer a stressful job?

A site reliability engineer (SRE) role can be stressful due to the responsibility of maintaining system uptime, handling incidents, and ensuring reliability under tight deadlines. The job often involves on-call duties, troubleshooting complex issues, and working with automation tools, which can contribute to work-related stress but also offers opportunities for skill development and problem-solving.

What is a site reliability engineer?

A site reliability engineer specializes in site reliability engineering, or SRE, a specific branch of operations first pioneered by Google. You are responsible for ensuring that when a website decides to scale a particular feature for various users to access, it does not break the underlying software or website functions. This means you need to use analytical problem-solving skills to determine how to make specific features on a new software release work on top of existing source code.

What are the key skills and qualifications needed to thrive as a site reliability engineer?

To thrive as a Site Reliability Engineer, you need a strong background in computer science, systems administration, and software engineering, often supported by a degree in a technical field. Familiarity with cloud platforms (like AWS or GCP), container orchestration (such as Kubernetes), infrastructure as code (Terraform or Ansible), and monitoring tools (Prometheus, Grafana) is typically expected. Strong problem-solving skills, effective communication, and a proactive mindset help SREs excel at incident management and cross-functional collaboration. These skills are crucial for maintaining system reliability, minimizing downtime, and driving continuous improvement in complex technical environments.

What are some of the most common challenges site reliability engineers face when balancing system reliability with rapid software delivery?

Site Reliability Engineers (SREs) often navigate the challenge of maintaining highly reliable systems while supporting fast-paced software releases. This involves managing incidents, automating processes to reduce manual toil, and working closely with development teams to embed reliability into the software development lifecycle. SREs must carefully prioritize their efforts between proactive improvements and urgent, reactive fire-fighting. Effective communication and collaboration with both operations and development teams are crucial to ensuring service uptime without slowing down innovation.

What is the difference between Site Reliability Engineer vs DevOps Engineer?

AspectSite Reliability EngineerDevOps Engineer
CredentialsTypically requires a computer science degree, certifications like AWS, Google Cloud, or KubernetesSimilar credentials, often with cloud certifications and scripting skills
Work EnvironmentFocuses on maintaining and improving system reliability, often in large-scale production environmentsWorks on automation, CI/CD pipelines, and deployment processes across development and operations teams
Industry UsageCommon in tech, cloud services, and large-scale enterprise companiesWidely used in software development, cloud, and IT organizations

Both roles require strong technical skills and cloud knowledge, but SREs focus more on system reliability and uptime, while DevOps engineers emphasize automation and deployment processes. They often collaborate but have distinct primary responsibilities.

What is a site reliability engineer?

A Site Reliability Engineer (SRE) is a professional who applies software engineering principles to infrastructure and operations problems. Their primary goal is to create scalable and highly reliable software systems, often bridging the gap between development and IT operations. SREs automate tasks, monitor system health, respond to incidents, and work to improve system reliability and performance. They also help define service level objectives (SLOs) and ensure systems meet customer expectations for uptime and availability.

What are the most commonly searched types of Site Reliability Engineer jobs in Sandy, UT?

The most popular types of Site Reliability Engineer jobs in Sandy, UT are:

What are popular job titles related to Site Reliability Engineer jobs in Sandy, UT?

For Site Reliability Engineer jobs in Sandy, UT, the most frequently searched job titles are:

What cities near Sandy, UT are hiring for Site Reliability Engineer jobs?

Cities near Sandy, UT with the most Site Reliability Engineer job openings:

Infographic showing various Site Reliability Engineer job openings in Sandy, UT as of August 2026, with employment types broken down into 1% As Needed, 81% Full Time, 15% Part Time, 1% Temporary, and 2% Contract. Highlights an 94% Physical, 3% Hybrid, and 3% Remote job distribution, with an average salary of $125,988 per year, or $60.6 per hour.

LINUX Administrator, Python, Production Support, SRE - CI/CD, 12+ Mths CTH South Jordan, UT

ZnA Inc

South Jordan, UT โ€ข On-site

$54.50 - $72.25/hr

Contractor

Re-posted 7 days ago


Job description

Production Support / SRE, LINUX/UNIX Administration, SQL, Automation ( Python/bash/ Perl/Ruby/Shell),  Messaging (MQ, CPS, XML, FIX), CI/CD ( Git, Artifactory, Jenkins, Docker), Data Streaming ( SPARK, Kafka),  observability and monitoring (Grafana, Splunk, Dynatrace, AppDynamics),  Configuration/Release mgt tools (Puppet, Ansible, Chef, GitHub) 12+ Mths Cont to Hire South Jordan, UT

JPC - 3551

Level 2: (6 to 8 yrs of total industry Exp)

Loc: South Jordan, UT (Hybrid: Hybrid - 3-days a week in office is mandatory)

Duration: 12+ months (Contract to Hire)

Interview Process:
- 1st: Zoom
- 2nd: Onsite
 

Description:


Bachelor's Degree: Yes
Industry Background: Plus
Years of experience: 2 - 5 years exp
Shifts:
• Morning: 8am - 5:00pm
• Evening: 12:30 - 8:am
o Weekend: OnCall (Remote)
Must:
Linux and Unix Hands on Experience!
o Needs to be comfortable troubleshooting at OS Level
Production Support Exp
o Incident Handling, Debugging live systems and on-call experience
Scripting/Automation experience within Python
o Not a Software Engineer but most be able to automate repetitive tasks
• Comm skills MUST BE THERE!
o Will be working hands on with Dev Teams and the business
• Service now Experience from a ticketing perspective
• Knowledge of ITIL Principles
Plus:
• Good understanding of Java, GO, C++, Scala etc
• Grafana and snowflake
• Any Cloud Exp
• Agentic AI background or knowledge
Description
Must haves:
Good hands on Linux experience and SQL
Knowledge for Python (more from a reading code perspective opposed from scripting)
This is a pipeline requisition to be used for all open consultant positions for Reliability & Production Engineering (RPE) in Utah. Please submit candidates who satisfy requirements for either Production Support or System Reliability Engineering (SRE) roles.
See below for detailed job descriptions:
1- PRODUCTION SUPPORT ANALYST
As Production Support Analyst, your responsibilities will include, but not be limited to:

  • Monitoring for and resolving issues across the entire tech stack: hardware, software, application and network. A majority of your time will be devoted to production support activities.
  • Working closely with engineering/development teams to address repetitive issues, reduce operational effort and the likelihood of future service disruptions.
  • Partnering with business users and other technology teams to manage significant events such as business continuity/disaster recovery tests, IPOs, stock splits, and major infrastructure changes.
  • Defining and refining standard operating procedures (SOP) for everything from monitoring to troubleshooting complex code and infrastructure issues.
  • Identifying and driving opportunities to improve platform supportability through automation.
  • Advocating for reliability priorities in application design reviews and operational readiness exercises for new and existing services.
  • Participating in weekend and off hours on-call rotation.
  • - Collaborating and striving to understand business users? needs and problems.
    Qualifications – External
    What skills and experience do I need? You should apply if you have at least a Bachelors degree in Computer Science or other technical discipline(s), plus hands-on experience with any combination of the following:
  • - 3-5+ years practical experience in production systems support or application development.
  •  Hands on experience managing systems in a large scale distributed Unix/Linux environment is essential.
  •  Effective communicator who is comfortable speaking in front of both internal/external groups as well as business clients-
  •  Demonstrated ability to troubleshoot problems and debug to conclusively identify root causes
  •  Knowledge of ITIL Principles. ITIL certification is a plus.
  •  Knowledge of Unix/Linux operating system level concepts such as processes, memory allocation, and networking, with an understanding of how applications are affected by these, and ability to debug and troubleshoot accordingly.
  • Automation-related experience is particularly valued, using scripting languages such as Python, bash, Perl, and/or Ruby. Higher-level compiled languages such as C++, C#, JAVA, Scala, and Go are a big plus.
  • - Working ability to interact with message transport platforms and protocols (MQ, CPS, XML, FIX) and distributed database technologies (DB2, Sybase, Mongo, GreenPlum, Postgres, KDB).
  • - Autosys scheduling and batch processing concepts- Experience with source code and binary repositories, build tools, and CI/CD (Git, Artifactory, Jenkins, Docker) etc and data streaming technologies like Spark, Kafka etc.
  • - Hands on experience on enterprise tools set such as Grafana, Splunk, Dynatrace, AppDynamics, etc.
  • Awareness of, and ability to reason through modern software & systems architectures, including load-balancing, queueing, caching, distributed systems failure modes generally, micro services, etc.
    2- SYSTEM RELIABILITY ANALYST
    As a System Reliability Analyst, your responsibilities will include, but not be limited to:
    - Working closely with engineering/development teams to design, build, optimize, and maintain systems.
    - Troubleshooting issues across the entire technology stack: hardware, software, application, and network.
    - Aggressively targeting toil and operational risk, and deploying solutions to reduce these.
    - Broadening infrastructure and application observability.
    - Proactively identifying and addressing active or potential risks to system reliability.
    - Advocating for reliability priorities in application design reviews and operational readiness exercises for new and existing services.
    Qualifications:
    - External What skills and experience do I need?
    You should apply if you have at least a Bachelor's degree in Computer Science or other technical discipline(s), plus hands-on experience with any combination of the following:
    - 3-5+ years practical experience in production systems support or application development- Hands on experience managing systems in a large scale distributed Unix/Linux environment is essential.
    - Automation-related experience is required, using scripting languages such as Python, bash, Perl, and/or Ruby. Higher-level compiled languages such as C++, C#, JAVA, Scala, and Go are a big plus.
    - Deep knowledge of and hands-on experience applying the principles of System/Site Reliability Engineering (SRE).
    - Practical experience designing and instrumenting SLO/SLI dashboards is particularly valuable.
    - Hands on experience on enterprise tools such as AppDynamics, Grafana, Splunk, Dynatrace
    - Experience with Puppet, Ansible, Chef, GitHub or any automation/configuration/release management tools- Awareness of, and ability to reason through modern software and systems architectures, including load
    -balancing, databases, queueing, caching, distributed systems failure modes, micro services, Cloud, etc.
    - Working ability to interact with message transport platforms and protocols (MQ, CPS, XML, FIX) and distributed database technologies (DB2, Sybase, Mongo, GreenPlum, Postgres, KDB).
    - Autosys scheduling and batch processing concepts.
    - Deep understanding of infrastructure and operating system concepts such as processes, memory allocation, and networking, with an understanding of how applications are affected by the above, and ability to debug and troubleshoot accordingly.