1

Ai Reliability Engineer Jobs in Chicago, IL (NOW HIRING)

LNVGY). We are looking for an AI Reliability Operations Engineer to support the operational health of Qira's production and non-production systems. Qira is Lenovo's cross-device Personal AI that ...

New

Reliability Engineer

Chicago, IL · Hybrid

$105K - $132K/yr

Reliability Engineer - IE08GE We're determined to make a difference and are proud to be an ... AI-Driven Automation: * Research and implement AI-based anomaly detection to predict infrastructure ...

Reliability Engineer

Chicago, IL · Hybrid

$105K - $132K/yr

Reliability Engineer - IE08GE We're determined to make a difference and are proud to be an ... AI-Driven Automation: * Research and implement AI-based anomaly detection to predict infrastructure ...

Sr. Site Reliability Engineer (SRE)

Chicago, IL · On-site

$58.75 - $78/hr

Moonlite AI delivers high-performance AI infrastructure for organizations running intensive computational research and large-scale model training. The Sr. Site Reliability Engineer will be ...

Define and lead WEX's AI-Powered Reliability Engineering strategy, driving adoption of SRE agents across the software lifecycle-from design and development through deployment and operations, to ...

Define and lead WEX's AI-Powered Reliability Engineering strategy, driving adoption of SRE agents across the software lifecycle-from design and development through deployment and operations, to ...

Site Reliability Engineer

Chicago, IL · On-site +1

$58.75 - $78/hr

The role combines traditional SRE responsibilities with modern AI-assisted engineering practices , leveraging AI tools to improve incident response, documentation, operational workflows, and ...

Driven by our investment in cutting-edge technologies like AI and cloud solutions, we're home to a ... Fitch Solutions SRE provides Service Reliability Engineering expertise to Fitch's development ...

Site Reliability Engineer

Riverwoods, IL · On-site

$59.25 - $78.75/hr

Are you interested in working with the World's leading AI-first Quality Engineering Company? Ready ... We are looking for a Site Reliability Engineer to join our growing team in Riverwoods, IL United ...

New

Staff Site Reliability Engineer (SRE) - Platform Engineering Note: This position follows a hybrid ... Strategic AI Integration: A mastery of leveraging Generative AI and Agentic workflows (e.g., Gemini ...

Staff Site Reliability Engineer (SRE) - Platform Engineering Note: This position follows a hybrid ... Strategic AI Integration: A mastery of leveraging Generative AI and Agentic workflows (e.g., Gemini ...

Staff Site Reliability Engineer (SRE) - Platform Engineering Note: This position follows a hybrid ... Strategic AI Integration: A mastery of leveraging Generative AI and Agentic workflows (e.g., Gemini ...

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Director of Cloud SRE

Downers Grove, IL · On-site +1

$57 - $75.50/hr

... Agentic AI deeper into our SRE ecosystem. While our primary application runtime is GCP, this leader must be equally comfortable partnering across the SRE organization to extend reliability and ...

New

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Site Reliability Engineer III

Chicago, IL

$58.75 - $78/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Banking ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Senior AI Quality & Reliability Engineer

Chicago, IL · On-site

$91K - $123K/yr

You will help advance Vizient's Quality Engineering capabilities beyond traditional software testing toward AI-native validation, AI-assisted testing, runtime observability, reliability engineering ...

next page

Showing results 1-20

Ai Reliability Engineer information

See Chicago, IL salary details

$62.8K

$121.5K

$145.3K

How much do ai reliability engineer jobs pay per year?

As of Aug 22, 2026, the average yearly pay for ai reliability engineer in Chicago, IL is $121,529.00, according to ZipRecruiter salary data. Most workers in this role earn between $105,600.00 and $132,900.00 per year, depending on experience, location, and employer.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Chicago, IL?

For Ai Reliability Engineer jobs in Chicago, IL, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Chicago, IL look for?

The top searched job categories for Ai Reliability Engineer jobs in Chicago, IL are:

What cities near Chicago, IL are hiring for Ai Reliability Engineer jobs?

Cities near Chicago, IL with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in Chicago, IL as of August 2026, with employment types broken down into 72% Full Time, 24% Part Time, and 4% Contract. Highlights an 65% Physical, 5% Hybrid, and 30% Remote job distribution, with an average salary of $121,529 per year, or $58.4 per hour.

AI Reliability Operations Engineer

Lenovo

Chicago, IL • On-site

$80 - $90/hr

Other

Posted yesterday

New


Lenovo rating

7.7

Company rating: 7.7 out of 10

Based on 19 frontline employees who took The Breakroom Quiz

80th of 159 rated electronics manufacturers


Job description

United States of America - Illinois - Chicago

Why Work at Lenovo

We are Lenovo. We do what we say. We own what we do. We WOW our customers.

Lenovo is a US$83 billion revenue global technology powerhouse, ranked #153 in the Fortune Global 500, and serving millions of customers every day in 180 markets. Focused on a bold vision to deliver Smarter Technology for All, Lenovo has built on its success as the world’s largest PC company with a full-stack portfolio of AI-enabled, AI-ready, and AI-optimized devices (PCs, workstations, smartphones, tablets), infrastructure (server, storage, edge, high performance computing and software defined infrastructure), software, solutions, and services. Lenovo’s continued investment in world-changing innovation is building a more equitable, trustworthy, and smarter future for everyone, everywhere. Lenovo is listed on the Hong Kong stock exchange under Lenovo Group Limited (HKSE: 992) (ADR: LNVGY).

We are looking for an AI Reliability Operations Engineer to support the operational health of Qira's production and non-production systems. Qira is Lenovo’s cross-device Personal AI that works across phones, PCs, and other Lenovo and Motorola products. This role spans system monitoring, alert response, incident response, and observability across the full AI stack, including model performance, inference pipelines, and cloud services. You will also have visibility into SDLC operations across staging and pre-production environments, helping ensure that releases and configuration changes land cleanly. This is a foundational role in keeping Qira stable and available for users around the world.

Location: Onsite in Chicago, IL (Hybrid, 3 days onsite, 2 days remote)

What You'll DoOperations Domain
  • Perform incident response: contain issues and work cross-functionally with dev teams to fully resolve them.
  • Monitor system health via Grafana dashboards, catching issues early and verifying resolution.
  • Serve as point of contact for change requests (CRs), triaging bug tickets from internal testers to the correct dev group.
  • Keep incident and CR records clear and accurate in ticketing systems.
Monitoring and Observability
  • Monitor production and non-production systems using observability dashboards, alerting tools, and AI-specific signals, including model performance, inference latency, and data pipeline health.
  • Watch proactively for early warning signals across Qira's cloud services, device integrations, and AI components, not just respond to alerts after they fire.
  • Review alert thresholds, update runbooks, and flag procedural gaps to the shift lead or SRE. This role is expected to improve the process, not just execute it.
Build Tooling and Documentation
  • Build and improve tooling and scripts to reduce manual, repetitive work across incident and CR handling.
  • Maintain documentation standards: keep runbooks, CR records, and process docs accurate and usable by anyone on the team.
Release Support
  • Observe and report on SDLC operations across staging and pre-production environments, flagging anomalies and supporting engineering teams during releases and configuration changes.
  • Verify system health before and after deployments and configuration changes, and assist engineering with deployment checks.
Basic Qualifications
  • Direct experience in incident response or an SRE-adjacent role, not just monitoring or support.
  • Experience with observability tools such as Grafana, Datadog, or cloud-native.
  • Experience with alerting tools such as PagerDuty or OpsGenie, and ticketing systems such as Jira or ServiceNow.
  • Experience Troubleshooting: can isolate where a problem actually lives (which service, which layer) without being handed the answer.
  • Experience in Azure, including core cloud concepts and how services in an Azure environment are monitored.
  • Ability to work assigned on-call coverage.
Preferred Qualifications
  • Experience in a technical operations, SRE, or production support environment.
  • Exposure to AI or ML systems, including awareness of how model quality and data pipelines are monitored.
  • Basic scripting ability or comfort reading and adapting existing scripts and runbook commands.
  • Experience working across time zones in a globally distributed team.
  • Working knowledge of SRE concepts: P50/P95/P99 latency, MTTA/MTTM/MTTR (MTTx), the four golden signals (latency, traffic, errors, saturation), and how to apply them to triage.
  • Clear, precise written communication in English, including accurate incident updates under pressure.
What Success Looks Like

A successful AI Reliability Operations Engineer detects issues early, responds to alerts quickly, performs accurate initial triage, and keeps clear and complete records across production and non-production environments. Their work directly supports system uptime and ensures that Qira delivers a reliable, consistent experience for users at all times

The base salary budgeted range for this position is $80K - $90K. Individuals may also be considered for bonus and/or commission.

Lenovo’s various benefits can be found on www.lenovobenefits.com.

We are an Equal Opportunity Employer and do not discriminate against any employee or applicant for employment because of race, color, sex, age, religion, sexual orientation, gender identity, national origin, status as a veteran, and basis of disability or any federal, state, or local protected class.

Additional Locations

United States of America - Illinois - Chicago

If you require an accommodation to complete this application, please contactability@lenovo.com

#J-18808-Ljbffr

What Lenovo employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom