1

Ai Reliability Engineer Jobs in Arizona (NOW HIRING)

GenAI Engineer - SRE

Phoenix, AZ ยท On-site

$56.50 - $75.25/hr

... , Platform Health, and Engineering Productivity initiatives ... The candidate will leverage Generative AI technologies, automation frameworks, and cloud-native ...

Site Reliability Engineer

Chandler, AZ ยท On-site

$56.25 - $74.50/hr

... ) Technical Skills 2 Technology|Open System|Shell scripting Technical Skills 3 Technology ... Our AI-driven solutions empower financial institutions to make smarter decisions, enhance customer ...

Site Reliability Engineer

Chandler, AZ ยท On-site

$110 - $140/hr

... SRE) Technical Skills 2 Foundational|Development process generic|Object Oriented Development ... Our AI-driven solutions empower financial institutions to make smarter decisions, enhance customer ...

New

Site Reliability Engineer

Chandler, AZ ยท On-site

$56.25 - $74.50/hr

... SRE) Technical Skills 2 Foundational|Development process generic|Object Oriented Development ... Our AI-driven solutions empower financial institutions to make smarter decisions, enhance customer ...

Site Reliability Engineer II

Chandler, AZ ยท On-site

$120 - $170/hr

# Site Reliability Engineer IIApex Systems, Inc.Chandler, AZContractJul 27, 2026Engineering## **Job ... By applying for this job, you agree to receive calls, AI-generated calls, text messages, or emails ...

New

Site Reliability Architect

Phoenix, AZ ยท On-site

$56.50 - $75.25/hr

Job details Job Role Engineering Consultant 3 Career Role Senior Project Manager - Engineering - US ... Our team leverages advanced technologies such as AI, ML, and predictive analytics to drive ...

Fox is hiring a Senior Engineer, AI Site Reliability to help build and operate infrastructure and platforms to support APIs around our live direct to consumer APIs for major live events such as the ...

Thermal Quality and Reliability Engineer

Phoenix, AZ ยท On-site

$101K - $128K/yr

The Role and Impact As a Packaging Quality and Reliability Engineer, you will play a critical role ... leadership for the AI era, enabling our customers to design leadership products, global ...

next page

Showing results 1-20

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What are popular job titles related to Ai Reliability Engineer jobs in Arizona? For Ai Reliability Engineer jobs in Arizona, the most frequently searched job titles are:
What cities in Arizona are hiring for Ai Reliability Engineer jobs? Cities in Arizona with the most Ai Reliability Engineer job openings:

AI-Enabled Platform/SRE Engineer

Trinite Consulting Group LLC

Phoenix, AZ โ€ข On-site

$55/hr

Other

Posted 10 days ago


Job description

Role: AI-Enabled Platform/SRE Engineer

Location:  Dallas, TX / Scottsdale, AZ (Hybrid)

Duration: Long Term Contract

Pay Rate: $55/hr on C2C
Note: 2nd round inperson 

Key Responsibilities

  • Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
  • Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
  • Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
  • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
  • Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacentre Kubernetes environments.
  • Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
  • Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.

Core Technical Skills

  • Site Reliability Engineering (SRE) โ€“ Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
  • Kubernetes Platform Engineering โ€“ 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
  • Cloud & Infrastructure Automation โ€“ Strong experience in Google Cloud Platform, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
  • Software Development โ€“ 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
  • Observability & Monitoring โ€“ Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.
  • API & Microservices Engineering โ€“ Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
  • AI-Driven Operations (AIOps) โ€“ Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.