1

Ai Reliability Engineer Jobs in Ohio (NOW HIRING)

Digital - Principal SRE (AI Engineer)

Columbus, OH · On-site +1

$53.50 - $71.25/hr

The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is responsible for ...

Digital - Principal SRE (AI Engineer)

Columbus, OH · On-site +1

$53.50 - $71.25/hr

The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is responsible for ...

Digital - Principal SRE (AI Engineer)

Columbus, OH · On-site +1

$55 - $73.25/hr

The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is responsible for ...

Digital - Principal SRE (AI Engineer)

Columbus, OH · On-site +1

$53.50 - $71.25/hr

The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is responsible for ...

Reliability Engineer

Columbus, OH · Hybrid

$91K - $136K/yr

Reliability Engineer - IE08GE We're determined to make a difference and are proud to be an ... AI-Driven Automation: * Research and implement AI-based anomaly detection to predict infrastructure ...

Reliability Engineer

Columbus, OH · Hybrid

$99K - $124K/yr

Reliability Engineer - IE08GE We're determined to make a difference and are proud to be an ... AI-Driven Automation: * Research and implement AI-based anomaly detection to predict infrastructure ...

Reliability Engineer

Columbus, OH · Hybrid

$91K - $136K/yr

Reliability Engineer - IE08GE We're determined to make a difference and are proud to be an ... AI-Driven Automation: * Research and implement AI-based anomaly detection to predict infrastructure ...

Reliability Engineer

Columbus, OH · Hybrid

$99K - $124K/yr

Reliability Engineer - IE08GE We're determined to make a difference and are proud to be an ... AI-Driven Automation: * Research and implement AI-based anomaly detection to predict infrastructure ...

Reliability Engineer - Vibration

Delaware, OH · On-site

$97K - $122K/yr

Why Join Us | DuPont Careers Reliability Engineer - Vibration DuPont is seeking a highly ... Experience with AI-enabled predictive maintenance technologies * Experience with Emerson AMS ...

Digital - Principal SRE

Columbus, OH · On-site +1

$53.50 - $71.25/hr

Description The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is ...

Digital - Principal SRE

Columbus, OH · On-site +1

$55 - $73.25/hr

Description The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is ...

Digital - Principal SRE

Columbus, OH · On-site +1

$53.50 - $71.25/hr

Description The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is ...

Digital - Principal SRE

Columbus, OH · On-site +1

$53.50 - $71.25/hr

Description The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is ...

$130 - $210/hr

About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools ... The Role As a Site Reliability Engineer (SRE) on the Cloud Platform team, you will shape the ...

next page

Showing results 1-20

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What job categories do people searching Ai Reliability Engineer jobs in Ohio look for?

The top searched job categories for Ai Reliability Engineer jobs in Ohio are:

What cities in Ohio are hiring for Ai Reliability Engineer jobs?

Cities in Ohio with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in Ohio as of August 2026, with employment types broken down into 81% Full Time, 17% Part Time, and 2% Contract. Highlights an 63% Physical, 4% Hybrid, and 33% Remote job distribution.

Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498)

84.51

Cincinnati, OH • On-site

Full-time

Posted 9 days ago


Job description

PLEASE NOTE:  This role is on the Kroger Technology & Digital (KTD) team.  KTD is the technology division of The Kroger Co. responsible for building, operating, and innovating the digital infrastructure and software that powers Kroger's retail stores, e-commerce platforms, supply chain, and corporate operations.  There is a strong collaboration between KTD and 84.51.

Role Summary 
The Senior Manager, AI Reliability Engineering at KTD leads the engineering discipline that makes enterprise AI operationally trustworthy at scale. As the enterprise moves from AI pilots to production systems that associates and customers depend on every day, this leader ensures that models, agents, copilots, AI gateways and shared runtime services are dependable, high-quality, responsive and cost-efficient by design.  This leader will stand up a new discipline from the ground up: defining what production-grade AI means, engineering the standards and automation that bake reliability and quality into every system, and shaping how build teams design for resilience from day one. The role blends engineering leadership, technical ownership and cross-functional influence. Reliability is a strategic enabler of adoption and velocity - the difference between experimenting with AI and confidently running it at scale. 


Key Responsibilities 
*Build the Discipline (0-to-1) 

*Define what production-grade, operationally trustworthy AI means for the enterprise, including standards and quality bars for availability, behavior, latency, cost, control and recovery. 
*Stand up the AI Reliability Engineering function, its charter, operating model, roadmap, talent model and engineering culture. 
*Position reliability as an enabler of AI adoption and velocity, creating the confidence that allows the business to scale AI responsibly and aggressively. 
*Engineer Reliability and Quality Into Systems 

*Partner with Platform, Model and Applied AI teams to embed resilience, testability, observability and safe failure modes into AI systems from architecture forward. 
*Build reliability tooling and automation, including self-healing, automated evaluations, quality-regression detection, guardrail instrumentation and safe deployment controls. 
*Establish service-level indicators, service-level objectives, error budgets and reliability scorecards that shape architecture, delivery and roadmap decisions. 
*Engineer for high availability, graceful degradation, capacity, disaster recovery and rapid restoration while reducing manual toil and systemic failure patterns. 
*Provide Production Readiness and Agent Onboarding

*Create production-readiness standards covering named business and engineering owners, support models, runbooks, telemetry, quality evaluations, security and Responsible AI controls, escalation paths, service objectives and lifecycle controls. 
*Lead launch-readiness reviews for new platforms, models and agents, and make evidence-based readiness decisions with clear exception and risk-acceptance paths. 
*Build a scalable onboarding model for both centrally developed agents and domain-owned agents operating on shared enterprise platforms. 
*Own Production Quality, Observability and AgentOps 

*Own live observability, production-quality signals and leadership visibility across model and agent behavior, drift, hallucination and quality rates, latency, tool failures and evaluations in production. 
*Partner with Responsible AI to translate offline evaluation, risk and safety standards into continuous, automated production signals and operational controls. 
*Own the operational capabilities of the Agent Control Center, including estate health, pause, isolation, rollback, shutdown and lifecycle controls for unsupported or persistently unreliable agents. 
*Drive Efficiency and Performance 

*Engineer for cost and performance at scale, optimizing inference cost, token efficiency, model and routing economics, tool usage and infrastructure consumption as first-class objectives. 
*Partner with value tracking, metering, product and finance teams to translate consumption and capacity signals into investment decisions and surface the reliability, quality, latency and cost tradeoffs that shape AI strategy. 
*Provide resilience and incident excellence 

*Establish agentic incident prevention and response practices, including severity frameworks, on-call, escalation, incident command, runbooks, communications and blameless learning. 
*Own detection, initial diagnosis, containment, recovery and coordinated L1/L2 engineering response; route L3 defects to the accountable platform, model or agent engineering team. 
Continuously improve detection, recovery, recurrence and deployment safety through automation, permanent engineered fixes, release gates, controlled rollout, automated rollback and recovery testing. 
*Lead and Influence 

*Hire, coach and grow a high-caliber multidisciplinary team spanning site reliability, observability, AgentOps, release engineering and production readiness, with a culture grounded in ownership, automation and blameless learning. 
*Operate horizontally across Platform Engineering, Applied AI, Models, Gateway, Infrastructure, Security and Responsible AI, elevating production engineering across all of them. 
*Act as a technical thought leader for production-grade AI and represent production reality in enterprise architecture, strategy and roadmap decisions. 
Required Qualifications: 
*Minimum 10+ years of experience in software, platform, cloud, SRE or production engineering, including 3 or more years leading engineering teams. 
*Track record of building engineering capabilities, standards, automation or platforms that measurably improved system quality and reliability at scale. 
*Strong technical depth in distributed systems, Kubernetes, cloud infrastructure on GCP and/or Azure, networking, identity, automation, CI/CD and modern observability engineering. 
*Experience establishing SLOs, error budgets, on-call practices, production-readiness reviews, incident management and post-incident improvement programs. 
*Direct experience with ML, AI or LLM systems in production, or demonstrated ability to master failure modes unique to models and agents, including drift, non-determinism, quality regression, prompt and context issues, and hallucination. 
*Ability to influence architecture and roadmaps, lead through ambiguity, partner across a matrixed organization and drive outcomes without direct authority. 
*Excellent technical and executive communication, with the ability to shape strategy with leaders and go deep with senior engineers. 
Preferred Qualifications: 
*Hands-on experience with LLMOps, model observability and evaluation tooling such as LangSmith, MLflow, Arize, Fiddler or custom evaluation pipelines. 
*Familiarity with agentic systems, including orchestration, tool use, MCP-based integrations, registries and their unique reliability and quality challenges. 
*Background bridging software or SRE practices with ML, data science and Responsible AI engineering cultures, including experience establishing a new engineering discipline. 
*Experience with multi-region, high-availability services, hybrid operating models and regulated or high-risk enterprise workloads. 
Measures of Success: 
*SLO attainment, service availability and production quality of AI platforms, models and agents. 
*Mean time to detect, contain and recover from incidents. 
*Production-onboarding time, production-readiness compliance and coverage of ownership, telemetry, runbooks, evaluations and shutdown controls. 
*Quality-regression detection and containment, change-failure, rollback and repeat-incident rates. 
*Reduction in operational toil, noisy alerts and manual interventions. 
*Cost, capacity, latency and token efficiency of production AI workloads. 
*Strength, engagement and growth of the AI Reliability Engineering team. 

#LI-SSS