1

Ai Reliability Engineer Jobs in Michigan (NOW HIRING)

AI Engineer

Dearborn, MI · On-site

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

AI Platform Engineer Location: Hybrid - 4 days/week onsite | Must be willing to interview in person ... Drive continuous improvements in reliability, scalability, and performance. Collaboration ...

Senior Platform Engineer

Ann Arbor, MI · On-site

$102K - $140K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Reliability: Define SLOs, improve observability, and lead incident response and root cause analysis ... AI-Enabled Platform Engineering: Leverage AI-assisted operations, GitHub, Copilot governance ...

Principal AI Engineer

Auburn Hills, MI

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

We are seeking a Principal AI Engineer with deep, hands‑on experience in Large Language Models ... Drive model evaluation, prompt optimization, and system reliability improvements based on ...

Senior Platform Engineer

Ann Arbor, MI · On-site

$102K - $140K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Reliability: Define SLOs, improve observability, and lead incident response and root cause analysis ... AI-Enabled Platform Engineering: Leverage AI-assisted operations, GitHub, Copilot governance ...

Senior Platform Engineer

Ann Arbor, MI

$102K - $140K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Reliability: Define SLOs, improve observability, and lead incident response and root cause analysis ... AI-Enabled Platform Engineering: Leverage AI-assisted operations, GitHub, Copilot governance ...

Senior Software Engineer, DevOps

Ann Arbor, MI · On-site +1

$160K - $190K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Utilidata is a fast-growing NVIDIA-backed AI company enabling AI data centers to dynamically ... Guide team members in DevOps practices, promoting a culture of reliability and excellence

The AI Engineer is responsible for building, deploying, and maintaining AI-powered applications and ... and reliability tradeoffs. * Support incident response for AI-related production issues and ...

The AI Engineer is responsible for building, deploying, and maintaining AI-powered applications and ... and reliability tradeoffs. * Support incident response for AI-related production issues and ...

AI Engineer

Grand Rapids, MI · On-site

$50K - $112K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

... AI Engineer, you will be at the forefront of transforming raw data into actionable insights ... reliability, and output groundedness - Optimizing open-weight language models, including LLaMA ...

AI Engineer

Detroit, MI · On-site

$50K - $112K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

... AI Engineer, you will be at the forefront of transforming raw data into actionable insights ... reliability, and output groundedness - Optimizing open-weight language models, including LLaMA ...

Partner with data governance, security, and infrastructure teams to ensure AI solutions meet enterprise standards for privacy, reliability, and compliance * Stay current on emerging AI engineering ...

Showing results 21-40

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Michigan?

For Ai Reliability Engineer jobs in Michigan, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Michigan look for?

The top searched job categories for Ai Reliability Engineer jobs in Michigan are:

What cities in Michigan are hiring for Ai Reliability Engineer jobs?

Cities in Michigan with the most Ai Reliability Engineer job openings:

Cloud Engineer - AWS Foundation Platform

Stellantis

Auburn Hills, MI • On-site

$52.75 - $70.50/hr

Full-time

Posted 25 days ago


Stellantis rating

7.5

Company rating: 7.5 out of 10

Based on 131 frontline employees who took The Breakroom Quiz

14th of 44 rated automakers


Job description

Stellantis is seeking a skilled and motivated Cloud Engineer to reinforce and support the design, implementation, and continuous improvement of our AWS Foundation Product - an AWS Landing Zone governing 1,500+ AWS accounts distributed across all Stellantis business functions and geographies (NAFTA, LATAM, Europe, EMEA, and APAC). As a Cloud Engineer, you will collaborate with cross-functional teams - including SecOps, FinOps, platform consumers, and global infrastructure teams - to build reusable platform components, drive DevOps maturity, and operate the AWS organizational layer as a product. AI tools are an integral part of daily engineering practice in this team.
Key Responsibilities:
  • Design and implement end-to-end automation of the AWS account lifecycle - provisioning, configuration, baselining, and decommissioning - achieving zero-touch operations at scale using Lambda, Step Functions, and EventBridge..
  • Build and maintain a self-service cloud platform (account vending machine, network factory, identity federation) enabling application teams to onboard and operate autonomously, backed by an Internal Developer Portal with AI-powered capabilities.
  • Own Infrastructure as Code (IaC) delivery pipelines across 1,500+ accounts using Terraform, enforcing GitOps practices, policy-as-code, drift detection, and automated testing at every stage.
  • Define and operate SLOs, Error Budgets, and an observability stack with AI-assisted anomaly detection across provisioning, cost, network, and security signals - enabling predictive operations rather than reactive firefighting.
  • Manage AWS Control Tower, Organizations, and Service Control Policies (SCPs), enforcing governance, least-privilege IAM, and compliance guardrails with automated self-healing when drift is detected.
  • Apply AI tools as a standard daily practice for IaC authoring, automation scripting, test generation, incident analysis, and runbook maintenance - maintaining full engineering ownership of all generated outputs.
  • Lead post-mortems and translate every incident into systemic engineering improvements; maintain runbooks and playbooks as versioned, executable artifacts.

Basic Qualifications:
  • Bachelor's degree in computer science, Information Technology, or related field.
  • Minimum 5 years of experience designing, deploying, and managing AWS cloud environments at enterprise scale in multi-account contexts.
  • Expertise in AWS multi-account services: AWS Control Tower, Organizations, Service Control Policies, and Landing Zone design patterns.
  • Expert-level Terraform and/or AWS CDK for Infrastructure as Code; GitOps delivery pipelines with drift detection and policy gates (GitHub Actions, GitLab CI...).
  • Proficiency in Python and/or Typescripts for building platform tooling, internal CLIs, automation services, and event-driven workflows (Lambda, Step Functions, EventBridge...).
  • Daily use of AI coding assistants (GitHub Copilot, Claude Code, Codex ...) with demonstrated ability to review, validate, and own AI-generated code at production quality.
  • Solid understanding of DevOps and SRE practices: DORA metrics, SLOs, error budgets, incident management, policy-as-code, and blameless post-mortem culture.
  • Strong knowledge of AWS networking: Transit Gateway, VPC design, Route 53, Direct Connect, and PrivateLink.
  • Experience with AWS security services: Security Hub, GuardDuty, Config Rules, IAM design, and automated compliance remediation.
  • AWS Certifications preferred: DevOps Engineer Professional, Solutions Architect Professional, or Security Specialty.
  • Familiarity with AIOps and AI/ML platform services (Amazon Bedrock, SageMaker, Amazon Q Business, DevOps Guru) is a strong plus.

What Stellantis employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom