1

Ai Reliability Engineer Jobs in Orem, UT (NOW HIRING)

Senior Site Reliability Engineer

Lehi, UT · On-site

$53.50 - $71/hr

Our AI-powered DigiCert ONE platform unifies PKI, DNS, and certificate lifecycle management, to ... By integrating reliability early, the SRE fosters a culture of shared responsibility while enabling ...

Senior Site Reliability Engineer

Lehi, UT · On-site

$53.50 - $71/hr

Our AI-powered DigiCert ONE platform unifies PKI, DNS, and certificate lifecycle management, to ... By integrating reliability early, the SRE fosters a culture of shared responsibility while enabling ...

AI DevOps Engineer

Sandy, UT · On-site

$50.25 - $68.75/hr

... , and Infrastructure teams to deliver end-to-end capabilities Improve system performance, reliability, and observability Key Responsibilities Design and develop scalable backend systems for AI ...

AI DevOps Engineer

Sandy, UT · On-site

$50.25 - $68.75/hr

... , and Infrastructure teams to deliver end-to-end capabilities Improve system performance, reliability, and observability Key Responsibilities Design and develop scalable backend systems for AI ...

New

AI DevOps Engineer

Sandy, UT · On-site

$50.25 - $68.75/hr

AI DevOps Engineer - Cloud AI Platforms At NICE, we are not just building software-we are ... , and Infrastructure teams to deliver end-to-end capabilities • Improve system performance ...

Sr. Observability Engineer

Lehi, UT · On-site

$98K - $134K/yr

... , or DevOps * 5+ years automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency * AI- and workflow-literate. You've used or built scripted and AI-assisted ...

You'll collaborate cross-functionally with engineering and operations teams to improve reliability ... AI may assist with things like writing s, scheduling interviews, or reviewing applications against ...

... and ongoing reliability. • Prototype rapidly, iterate with live interaction data, and ... engineers, product managers, and AI/ML scientists to deliver end-to-end features that power ...

Staff AI Engineer The Role We are looking for a talented Staff AI Engineer to help build and scale ... Own the full lifecycle: architecture, implementation, deployment, and ongoing reliability.

Staff AI Engineer The Role We are looking for a talented Staff AI Engineer to help build and scale ... Own the full lifecycle: architecture, implementation, deployment, and ongoing reliability.

Senior AI Engineer - Agentic

Lehi, UT · On-site +1

$98K - $134K/yr

Own the full lifecycle: architecture, implementation, deployment, and ongoing reliability ... Collaborate with engineers, product managers, and AI/ML scientists to deliver end-to-end features ...

Visa Sponsorship: This AI Engineer position is currently not eligible for employment visa ... reliability. * Design, debug, test, and deploy software using Python and/or Java fundamentals ...

Visa Sponsorship: This AI Engineer position is currently not eligible for employment visa ... reliability. * Design, debug, test, and deploy software using Python and/or Java fundamentals ...

next page

Showing results 1-20

Ai Reliability Engineer information

See Orem, UT salary details

$53K

$102.6K

$122.6K

How much do ai reliability engineer jobs pay per year?

As of Aug 17, 2026, the average yearly pay for ai reliability engineer in Orem, UT is $102,562.00, according to ZipRecruiter salary data. Most workers in this role earn between $89,100.00 and $112,100.00 per year, depending on experience, location, and employer.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Orem, UT?

For Ai Reliability Engineer jobs in Orem, UT, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Orem, UT look for?

The top searched job categories for Ai Reliability Engineer jobs in Orem, UT are:

What cities near Orem, UT are hiring for Ai Reliability Engineer jobs?

Cities near Orem, UT with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in Orem, UT as of August 2026, with employment types broken down into 72% Full Time, 24% Part Time, and 4% Contract. Highlights an 66% Physical, 4% Hybrid, and 30% Remote job distribution, with an average salary of $102,562 per year, or $49.3 per hour.

Senior Site Reliability Engineer

DigiCert

Lehi, UT • On-site

$53.50 - $71/hr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

Re-posted 25 days ago


DigiCert rating

7.9

Company rating: 7.9 out of 10

Based on 5 frontline employees who took The Breakroom Quiz

79th of 224 rated it services


Job description

Who we are
DigiCert is a global leader in intelligent trust. We protect the digital world by ensuring the security, privacy, and authenticity of every interaction. Our AI-powered DigiCert ONE platform unifies PKI, DNS, and certificate lifecycle management, to secure infrastructure, software, devices, messages, AI content and agents. Learn why more than 100,000 organizations, including 90% of the Fortune 500, choose DigiCert to stop today's threats and prepare for a quantum-safe future at www.digicert.com
Job summary
The Site Reliability Engineer (SRE) collaborates with development teams to embed reliability, scalability, and performance best practices throughout the software development lifecycle. This role bridges software engineering and cloud operations, ensuring mission-critical systems remain highly available and resilient. By integrating reliability early, the SRE fosters a culture of shared responsibility while enabling rapid and safe feature delivery.
What you will do
  • Design and build fault-tolerant, high-performing systems that meet Service Level Objectives (SLOs) and Service Level Agreements (SLAs).
  • Implement monitoring, alerting, distributed tracing, and logging to ensure real-time system health visibility and proactive issue resolution.
  • Act as a first responder for production incidents, conduct blameless postmortems, and drive root cause analysis (RCA) and corrective actions.
  • Develop self-healing, automated deployments, and scaling solutions to minimize toil and improve system efficiency.
  • Improve continuous integration and deployment pipelines to enable safe, rapid, and reliable feature rollouts.
  • Review code, debug issues, and perform quality assurance (QA) on software components to enhance system reliability and performance.
  • Work closely with development teams to ensure best practices in software architecture, coding standards, and operational readiness.
  • Forecast scalability needs and optimize cloud infrastructure costs while balancing performance and efficiency.
  • Ensure production environments meet security and compliance requirements, collaborating with teams to mitigate vulnerabilities and enforce best practices.
  • Work closely with development teams to embed reliability at every stage rather than treating it as an afterthought.
  • Use error budgets to balance feature velocity with system stability.
  • Implement observability and automation-first principles to measure system health and drive continuous improvement.
  • Leverage game days, chaos engineering, and resilience testing to validate system robustness and refine operational processes.

What you will have
  • Extensive experience in distributed systems, cloud-native architectures (AWS, GCP, Azure), and DevOps practices.
  • Proficiency in Kubernetes, Terraform, CI/CD pipelines, and Infrastructure as Code (IaC).
  • Strong scripting and automation skills in Python, Go, Bash, or similar languages.
  • Expertise in observability tools such as Prometheus, Grafana, Datadog, Splunk, New Relic, and OpenTelemetry.
  • Ability to troubleshoot complex production issues and drive scalable, resilient solutions.
  • Experience reviewing code, debugging applications, and conducting software testing to ensure high reliability and quality.

Benefits
  • Competitive compensation and comprehensive health, dental, and vision coverage
  • Retirement savings programs with company matching (401(k) or RRSP)
  • Generous paid time off, including holidays, and vacation
  • Paid parental leave and family support benefits
  • Life and disability coverage
  • Flexible spending and health savings options (where applicable)
  • Health and wellness support, including gym reimbursement and wellness programs
  • Employee Assistance Program with 24/7confidential support for employees and families
  • Education assistance and professional development opportunities
  • Access to LinkedIn Learning and continuous learning resources
  • Employee referral bonus program and additional company perks and discounts
  • Internal rewards and recognition platform (Motivosity) to celebrate and acknowledge project wins, milestone achievements, and the outstanding contributions of our colleagues
  • Business travel insurance and global employee support programs

To protect candidate information and maintain a secure hiring process, all applications must be submitted through our careers portal. Resumes or CVs sent directly via email will not be reviewed or considered.
DigiCert is an Equal Opportunity employer and is committed to diversity in its workforce. In compliance with applicable federal and state laws, DigiCert prohibits discrimination on the basis of race or ethnicity, religion, color, national origin, sex, age, sexual orientation, gender identity/expression, veteran's status, status as a qualified person with a disability, or genetic information. Individuals from historically underrepresented groups, such as minorities, women, qualified person with disabilities, and protected veterans are strongly encouraged to apply.
#LI-RR1
Compensation Transparency:
The annualized base salary range for this position is outlined below.
Each candidate's compensation offer will be determined based on factors including experience, skills, qualifications, job duties, business needs, and location. For roles that include additional compensation components, total compensation may include base pay, bonus, equity, or other incentives.
This role may also be eligible for benefits, which will be discussed during the hiring process. We are committed to fair and transparent pay practices and comply with all applicable pay transparency requirements. If you would like more information about compensation or benefits, we are happy to provide additional details during the hiring process.
For more information regarding our comprehensive benefits, see the benefits section.
Base Salary
$125,000-$145,000 USD

What DigiCert employees say

Hours and flexibility

Workplace

Get the full story on Breakroom