1

Ai Reliability Engineer Jobs in Ohio (NOW HIRING)

$103.60 - $172.67/hr

As Team Lead Site Reliability Engineering (all genders), you own the reliability, scalability, and ... You make AI a core part of our operations: diagnosis, automation, monitoring, and insight. Your ...

Lead AI Engineer

Columbus, OH · On-site

$96K - $181K/yr

The Lead AI Engineer partners closely with product, platform, voice, and operations teams to ... Apply SRE principles to ensure resiliency, scalability, monitoring, and production readiness.

About this role: Wells Fargo is seeking a Lead Site Reliability Engineer (SRE) to support our ... Experience leveraging AI-powered engineering tools and platforms (e.g., GitHub Copilot, ChatGPT ...

AI Agent Engineer

Dayton, OH · On-site

$76K - $102K/yr

The AI Agent Engineer plays a critical role in ensuring AI agents run safely, reliably, and ... Monitor agent behavior, outputs, and reliability in production environments, identifying ...

AI Agent Engineer

Dayton, OH · On-site

$76K - $102K/yr

The AI Agent Engineer plays a critical role in ensuring AI agents run safely, reliably, and ... Monitor agent behavior, outputs, and reliability in production environments, identifying ...

$148K - $249K/yr

Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI. With a world-class ... Qualifications: - 5+ years software engineering or systems/performance engineering experience (BS ...

Senior Software Engineer

Columbus, OH · On-site

$118K - $156K/yr

This role is responsible for ensuring the reliability, availability, and performance of critical ... In addition, you will help lead the adoption of Wells Fargo's AI-enabled tools and capabilities ...

$102.72 - $171.19/hr

... combining AI methods with industrially proven optimization methods. The innovative industry ... Actively drive the SaaS transformation of our business unit while establishing modern SRE practices ...

Sr Lead Infrastructure Engineer

Columbus, OH · On-site

$102K - $138K/yr

Supports SRE teams as needed through technical consultation and escalation support for complex incidents and problem management, driving long-term fixes over one-off remediation * Leverages AI ...

VP Cloud Platform Engineering

Columbus, OH

$173K - $224K/yr

Experience leading enterprise cloud transformation, platform product models, FinOps, SRE, DevSecOps, and AI-enabled engineering initiatives. * Experience within highly regulated industries, including ...

Showing results 41-60

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What are popular job titles related to Ai Reliability Engineer jobs in Ohio? For Ai Reliability Engineer jobs in Ohio, the most frequently searched job titles are:
What job categories do people searching Ai Reliability Engineer jobs in Ohio look for? The top searched job categories for Ai Reliability Engineer jobs in Ohio are:
What cities in Ohio are hiring for Ai Reliability Engineer jobs? Cities in Ohio with the most Ai Reliability Engineer job openings:
Infographic showing various Ai Reliability Engineer job openings in Ohio as of August 2026, with employment types broken down into 76% Full Time, and 24% Contract. Highlights an 86% In-person, and 14% Remote job distribution.

Team Lead - Site Reliability Engineering (all genders)

FACT-Finder

On-site

$103.60 - $172.67/hr

Other

Posted 6 days ago


Job description

Introduction

FACT-Finder builds product discovery technology for eCommerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity. We are actively modernizing our hosting toward Kubernetes on Harvester – as an on-prem hybrid with the option to scale fully into the cloud in the mid-term. As Team Lead Site Reliability Engineering (all genders), you own the reliability, scalability, and cost of our hosting environments, drive this transformation end-to-end, and lead the team that delivers it.

Your mission
  • You own the operational health of our hosting across on-premise (Frankfurt, Stockholm) and cloud – availability, performance, and incident management.
  • You actively drive the modernization toward Kubernetes on Harvester: cluster topology, storage (Longhorn), networking (VLAN, load balancing, ingress), backup, and disaster recovery.
  • You build a production-grade k8s platform: lifecycle, upgrades, RBAC, secrets, GitOps (Argo CD / Flux), observability, and policy guardrails.
  • You shape the NG Search Operator (custom Kubernetes operator) and solve auto-scaling (HPA, VPA, KEDA, cluster autoscaler) for the current architecture.
  • You concretely define our on-prem hybrid model: which workloads run where, how we burst into the cloud, how we keep latency and cost under control – while keeping the architecture portable enough for a future cloud-only move.
  • You own capacity planning and hosting cost and turn cost into a deliberate, managed lever.
  • You lead and develop our currently 4-person Hosting team, own performance and technical direction, and set the standards and ownership culture.
  • You make AI a core part of our operations: diagnosis, automation, monitoring, and insight.
Your profile
  • Strong background in infrastructure or platform engineering across on-premise and cloud.
  • Hands-on depth with Kubernetes in production: cluster lifecycle, upgrades, networking, storage, RBAC, observability, GitOps delivery.
  • Proven people leadership experience, excellent communication and stakeholder management skills.
  • Ideally practical experience with Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack).
  • Experience leading a real migration from bare metal / classic VMs to a k8s-based platform – including stateful workloads, storage migration, cutover, and rollback.
  • Comfort designing or operating Kubernetes operators (custom controllers / CRDs), ideally for stateful systems like search, databases, or streaming.
  • Solid grasp of auto-scaling primitives (HPA, VPA, cluster autoscaler, KEDA) and how they interact with capacity planning on-prem and in the cloud.
  • Experience with on-prem hybrid architectures and owning reliability, capacity, and cost for production systems.
  • Hands-on fluency with AI tools in day-to-day operations.
  • Fluent English; German is a plus.
THE JOY OF WORKING WITH US
  • Impact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.
  • Leadership with real scope: You lead an established team and shape our platform in a decisive phase of our transformation.
  • Modern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud – with room to build things right.
  • AI-first mindset: We use AI as a real part of our daily work, not as a buzzword.
  • Ownership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.
  • Flexible work: Hybrid work model with a focus on outcomes.
  • Strong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.
  • Attractive benefits: Competitive salary, modern equipment, learning budget, and regular team events.
Job Location

Berlin, Munich, Pforzheim or Stockholm (hybrid)

#J-18808-Ljbffr