1

Ai Reliability Engineer Jobs in Highlands Ranch, CO

Principal Site Reliability Engineer

Denver, CO · On-site

$58.75 - $78/hr

We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data ... At Vertafore, we view reliability as a core engineering responsibility. You will operate ...

Sr. Site Reliability Engineer

Denver, CO · Hybrid

$58.75 - $78/hr

We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data ... Reliability is a core engineering responsibility, requiring strong software engineering skills and ...

Sr. Site Reliability Engineer

Denver, CO · On-site

$58.75 - $78/hr

We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data ... Reliability is a core engineering responsibility, requiring strong software engineering skills and ...

Sr. Site Reliability Engineer

Denver, CO · On-site

$58.75 - $78/hr

We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data ... Reliability is a core engineering responsibility, requiring strong software engineering skills and ...

Sr. Site Reliability Engineer

Denver, CO · Hybrid

$58.75 - $78/hr

We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data ... Reliability is a core engineering responsibility, requiring strong software engineering skills and ...

Staff SRE - Observability

Denver, CO · On-site

$58.75 - $78/hr

Experience implementing SRE practices including error budgets and toil metrics * Proficiency in ... Experience with AI/ML model deployment and monitoring in production environments Leadership ...

Director, Site Reliability Engineering

Denver, CO · On-site

$58.75 - $78/hr

We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data ... (SRE) will lead reliability, performance, and observability initiatives for a portfolio of ...

Director, Site Reliability Engineering

Denver, CO · On-site

$58.75 - $78/hr

We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data ... (SRE) will lead reliability, performance, and observability initiatives for a portfolio of ...

Director, Site Reliability Engineering

Denver, CO · On-site

$58.75 - $78/hr

We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data ... (SRE) will lead reliability, performance, and observability initiatives for a portfolio of ...

Director, Site Reliability Engineering

Denver, CO · On-site

$58.75 - $78/hr

We equip agencies, MGAs, and carriers with the core digital systems, specialized AI, and data ... (SRE) will lead reliability, performance, and observability initiatives for a portfolio of ...

Site Reliability Engineer II

Denver, CO · On-site +1

$98K - $138K/yr

The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining ... We may use artificial intelligence (AI) tools to support parts of the hiring process, such as ...

Site Reliability Engineer II

Denver, CO · On-site +1

$98K - $138K/yr

The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining ... We may use artificial intelligence (AI) tools to support parts of the hiring process, such as ...

Showing results 21-40

Ai Reliability Engineer information

See Highlands Ranch, CO salary details

$64K

$123.8K

$148K

How much do ai reliability engineer jobs pay per year?

As of Aug 20, 2026, the average yearly pay for ai reliability engineer in Highlands Ranch, CO is $123,826.00, according to ZipRecruiter salary data. Most workers in this role earn between $107,600.00 and $135,400.00 per year, depending on experience, location, and employer.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Highlands Ranch, CO?

For Ai Reliability Engineer jobs in Highlands Ranch, CO, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Highlands Ranch, CO look for?

The top searched job categories for Ai Reliability Engineer jobs in Highlands Ranch, CO are:

What cities near Highlands Ranch, CO are hiring for Ai Reliability Engineer jobs?

Cities near Highlands Ranch, CO with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in Highlands Ranch, CO as of August 2026, with employment types broken down into 100% Full Time. Highlights an 67% In-person, and 33% Remote job distribution, with an average salary of $123,826 per year, or $59.5 per hour.

Senior Site Reliability Engineer

Akamai Technologies, Inc.

Denver, CO • On-site

$121.40 - $218.60/hr

Other

Medical, Life, Retirement, PTO

This job post has expired today. Applications are no longer accepted.


Akamai Technologies rating

8.0

Company rating: 8.0 out of 10

Based on 12 frontline employees who took The Breakroom Quiz

123rd of 245 rated software companies


Job description

Do you enjoy collaborating with teams to solve complex challenges?

Do you enjoy solving large scale distributed content delivery challenges?

Join our critical AI Hardware SRE Team!

The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be responsible for ensuring best-in-class uptime and reliability of our AI hardware infrastructure offerings.

Partner with the best

In this role, you'll play a part in pioneering the reliability an elite, high-density hardware and software infrastructure spanning the globe. You'll collaborate with product teams from the earliest stages of development to ensure the reliability, scalability, and performance of our systems. You'll define key performance indicators and defend them when they are breached.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Developing and scaling robust programmatic tooling and infrastructure-as-code utilities in Python to eliminate operational toil and automate fleet-wide provisioning.

  • Integrating automated workflows across disconnected corporate ticketing systems to optimize time-to-mitigate metrics for hardware and network break-fix events.

  • Leveraging advanced AI utilities and LLM-assisted development paradigms where appropriate to accelerate technical execution, script authorship, and system analysis.

  • Working on cutting-edge private cloud and compute technologies to improve the availability, latency, and overall systemic health of high-density hardware environments.

  • Designing and implementing telemetry pipelines, custom Prometheus/Grafana monitoring dashboards, and AI-based anomaly detection tailored for bare-metal and virtualized environments.

  • Participating in 24x7x365 on-call rotations, spearheading real-time incident management, and managing high-severity service disruption protocols via automated PagerDuty and Slack workflows.

  • Partnering directly with third-party infrastructure vendors and coordinating on-site field technicians to facilitate uptime activities.

Do what you love

To be successful in this role you will:

  • Have 5 years of relevant experience and a Bachelor's degree in Computer Engineering, Computer Science or equivalent.

  • Possess tooling and coding ability in languages like Python to construct scalable operational tools, API integrations, and automation frameworks.

  • Show hands-on experience with modern observability stacks and timeseries engines, like Prometheus, Grafana, OpenTelemetry, and Loki.

  • Possess a working understanding of advanced networking topologies, high-bandwidth routing/switching infrastructure, BGP, and dual-stack IPv4/IPv6 networks.

  • Have experience acting as a key designer for new service rollouts, including establishing operational readiness criteria, telemetry baselines, and alerting thresholds.

  • Demonstrate extensive experience building technical runbooks, leading complex incident response bridges, and driving comprehensive, blameless post-mortems.

  • Display a proven ability to take absolute ownership of ambiguous technical problems, coordinate cross-functional teams, and drive for production-grade solutions.

About us

At Akamai, we make life better for billions of people, trillions of times a day.

Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.

Our focus is simple:

Cloud and Edge: Running apps closer to users for instant performance.

Security: Neutralizing threats before they ever reach your data.

Content Delivery: Scaling the world's biggest moments without a glitch.

AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.

At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai

We support your health, well-being, finances, and life beyond work. We offer competitive benefits, including healthcare, 401K savings plan, company holidays, PTO, sick time, family-friendly benefits such as parental leave and an employee assistance program focusing on mental and financial wellness.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.

We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Compensation

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $121,400 - $218,600/year; a candidate’s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP).

Equal Employment Opportunity Rights

Akamai Technologies is an affirmative action, Equal Opportunity Employer that values the strength that diversity brings to the workplace. All qualified applicants will receive consideration for employment and will not be discriminated against on the basis of gender, gender identity, sexual orientation, race/ethnicity, protected veteran status, disability, or other protected group status.

#J-18808-Ljbffr

What Akamai Technologies employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom