1

Ai Reliability Engineer Jobs in Missouri (NOW HIRING)

Design, build, and maintain highly available, scalable, and secure infrastructure to support our AI ... Core Engineering: 10+ years of experience in SRE, DevOps, or Software/Systems Engineering ...

New

$44 - $58.50/hr

Our partner is looking for a Site Reliability Engineer based in Netherlands. Join a technology ... We use an AI-powered matching process to ensure your application is reviewed quickly, objectively ...

Software Engineer 2 (DevOps/SRE)

Earth City, MO · On-site

$54 - $71.75/hr

Leverage AI/ML-enabled capabilities, automation, and advanced analytics where applicable to improve ... Embed software reliability engineering (SRE) and observability practices into development workflows

$44 - $58.50/hr

This is an opportunity to shape the reliability of next-generation AI infrastructure within an innovative and globally distributed environment. Accountabilities: As an AI Observability Engineer, you ...

New

Role Specific Information About the Role As Senior Reliability Engineer, you will ensure the ... Passion for and experience with AI and ML methodologies (MLOps) * Experience writing Infrastructure ...

Our partner is looking for an AI Platform Engineer based in Netherlands. This role offers the ... AI reliability are considered a strong advantage. Benefits * Fully remote opportunity with the ...

next page

Showing results 1-20

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What are popular job titles related to Ai Reliability Engineer jobs in Missouri? For Ai Reliability Engineer jobs in Missouri, the most frequently searched job titles are:
What job categories do people searching Ai Reliability Engineer jobs in Missouri look for? The top searched job categories for Ai Reliability Engineer jobs in Missouri are:
What cities in Missouri are hiring for Ai Reliability Engineer jobs? Cities in Missouri with the most Ai Reliability Engineer job openings:

Staff Site Reliability Engineer

Jobtailor

California, MO • On-site

$140 - $210/hr

Other

Posted 3 days ago

New


Job description

  • System Resilience: Design, build, and maintain highly available, scalable, and secure infrastructure to support our AI-native cybersecurity platform.
  • Automation & Tooling: Develop internal tooling and automation to streamline deployment processes, incident response, and capacity planning.
  • Performance Engineering: Monitor system performance and proactively identify bottlenecks, optimizing infrastructure for low-latency, high-throughput AI workloads.
  • Incident Management: Lead incident response efforts, conduct post-mortems, and implement long-term solutions to prevent recurring reliability issues.
  • Infrastructure as Code (IaC): Manage infrastructure via code, driving consistency, auditability, and scalability across our cloud environments (e.g., AWS, GCP).
  • Cross-Functional Collaboration: Partner with sibling Engineering teams, Product, and Security teams to ensure reliability is baked into our development lifecycle from concept to production.
Requirements
  • Core Engineering: 10+ years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing production systems at scale.
  • Cloud Infrastructure: Deep expertise in public cloud environments (AWS, GCP, or Azure) and managing services such as Kubernetes (EKS/GKE), networking, and storage.
  • Infrastructure as Code: Extensive experience with tools like Terraform, Pulumi, or similar technologies to manage complex infrastructure deployments.
  • Observability: Hands-on experience with monitoring, logging, and tracing stacks (e.g., Prometheus, Grafana, ELK, Datadog) to drive data-informed reliability decisions.
  • Distributed Systems: Solid understanding of microservices architecture, distributed databases, and event-driven systems.
  • Communication: Clear, concise communication skills and a bias for collaborative problem-solving.
  • Leadership Alignment: Proven track record of guiding multi-stakeholder initiatives and influencing engineering practices across teams.
  • Analytical Rigor: Strong problem-solving, debugging, and analytical skills, especially in high-pressure environments.
  • Domain Background: Prior work in cybersecurity, specifically regarding SIEM, EDR, or SOAR infrastructure is nice-to-have.
  • AI/ML Infrastructure: Experience supporting infrastructure for large-scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization) is nice-to-have.
  • Startup Mentality: Background driving high-impact engineering initiatives in high-growth startups or enterprise SaaS is nice-to-have.
  • Strong familiarity with Agentic Workflows such as Agno, Temporal, etc. is nice-to-have.
Core Competencies

Demonstrates expertise in designing and maintaining scalable, secure cloud infrastructure, with a strong focus on automation, performance engineering, and incident management. Proven ability to collaborate across teams and drive reliability in AI-native cybersecurity platforms.

Highest-signal resume keywords
  • 10+ Years Experience in SRE, DevOps, or Software Engineering
  • Deep Expertise in AWS, GCP, or Azure
  • Extensive Experience with Terraform or Pulumi
  • Hands-On Experience with Prometheus, Grafana, or ELK
  • Strong Problem-Solving and Analytical Skills
ATS Optimization Keywords Hard Skills
  • Infrastructure as Code
  • Performance Engineering
  • Incident Management
  • Cloud Infrastructure Management
  • Distributed Systems Understanding
  • AI/ML Infrastructure Support
  • Automation Development
  • Monitoring and Observability
  • Capacity Planning
  • Microservices Architecture
Soft Skills
  • Clear Communication Skills
  • Collaborative Problem-Solving
  • Leadership in Multi-Stakeholder Initiatives
Industry Keywords
  • Cybersecurity
  • AI-Native Platforms
  • High-Impact Engineering
  • Startup Mentality
  • Enterprise SaaS
Tools & Technologies
  • Kubernetes (EKS/GKE)
  • Prometheus
  • Grafana
  • ELK
  • Datadog
  • Terraform
  • Pulumi
  • Agentic Workflows
  • SIEM
  • EDR
#J-18808-Ljbffr