1

Ai Reliability Engineer Jobs in California (NOW HIRING)

You'll also help define how WEX applies AI to reliability engineering, building agents and reusable skills that automate high-TOIL work, integrating safely into our AI ecosystem, and establishing ...

You'll also help define how WEX applies AI to reliability engineering, building agents and reusable skills that automate high-TOIL work, integrating safely into our AI ecosystem, and establishing ...

Reliability Engineer

Costa Mesa, CA · On-site

$146K - $194K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

Staff Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

You'll also help define how WEX applies AI to reliability engineering, building agents and reusable skills that automate high-TOIL work, integrating safely into our AI ecosystem, and establishing ...

We're a family-founded company on a mission to create the world's first AI-powered Personal ... The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the ...

Reliability Engineer

Costa Mesa, CA · On-site

$146K - $194K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

Reliability Engineer

Costa Mesa, CA · On-site

$110K - $138K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

Showing results 21-40

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What job categories do people searching Ai Reliability Engineer jobs in California look for?

The top searched job categories for Ai Reliability Engineer jobs in California are:

What cities in California are hiring for Ai Reliability Engineer jobs?

Cities in California with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in California as of August 2026, with employment types broken down into 100% Full Time. Highlights an 60% In-person, and 40% Remote job distribution.

Site Reliability Engineering (SRE) Manager, Apple Maps

Apple

Cupertino, CA

$267K - $401K/yr

Full-time

Medical, Dental, Retirement

Re-posted 23 days ago


Key responsibilities

  • Operate with the fundamental principle that reliability is feature number one

  • Define and drive the strategic roadmap for Apple Maps serving infrastructure in partnership with engineering and product teams

  • Lead and grow an SRE organization, setting the bar for technical excellence, operational rigor, and engineering culture


Apple rating

8.1

Company rating: 8.1 out of 10

Based on 684 frontline employees who took The Breakroom Quiz

6th of 30 rated technology retailers


Job description

Apple Maps and location services are used by hundreds of millions of people every day to navigate the world. Behind every search, route, and point of interest is a massive distributed serving infrastructure that must be fast, reliable, and always available.
This SRE org is responsible for the availability and automation of some of the most visible and widely used services that power Apple Maps.
If you are passionate about reliability at planet scale and are excited to help us build, grow, manage and deliver infrastructure that scales with Apple Maps, this is the opportunity for you!
Description
We are looking for a senior SRE leader to set the strategic direction for our Maps serving infrastructure. This is not just a "keep the lights on" role - this is a leadership position that defines where our infrastructure goes next, how our SRE practice evolves, and how we build the teams and partnerships to get there. You will work closely with engineering, product, and operations partners across Apple to shape the roadmap for our serving platform. You will lead an organization of SREs and hold responsibility for the reliability, scalability, and operational excellence of some of Apple's most visible services. We believe AI will fundamentally reshape how SRE is practiced - from incident detection and resolution to capacity planning and toil elimination - and we're looking for a leader who shares that conviction and can drive that transformation across the organization.","responsibilities":"Operate with the fundamental principle that reliability is feature number one
Define and drive the strategic roadmap for Apple Maps serving infrastructure in partnership with engineering and product teams
Lead and grow an SRE organization, setting the bar for technical excellence, operational rigor, and engineering culture
Represent the SRE perspective in cross-functional planning - translating reliability requirements into architecture decisions and investment priorities
Champion the "Engineering" in Site Reliability Engineering - driving automation, platform improvements, capacity strategy, and systems design beyond reactive incident response
Define and execute a clear AI strategy for the SRE organization - identifying high-impact opportunities where AI/ML-powered tooling can reduce toil, accelerate root cause analysis, and improve reliability outcomes
Drive adoption of AI-assisted tooling (copilots, intelligent runbooks, LLM-based diagnostics, anomaly detection) into day-to-day SRE workflows
Build a culture where engineers actively experiment with AI tools and modern approaches to solve operational problems
Build strong partnerships across Apple, negotiating priorities and aligning on shared goals with an Apple-first mindset
Communicate effectively at the executive level - presenting strategy, trade-offs, and progress to senior leadership
Mentor and develop leaders within your organization, creating a culture where people do their best work
Celebrate wins, recognize contributions, and make tough calls when needed
Preferred Qualifications
Experience with cloud infrastructure (AWS, GCP) and Kubernetes at scale
Background in capacity planning, performance engineering, or infrastructure architecture
Track record of driving cultural and process transformation within SRE organizations
Experience building or deploying AI-powered operational tooling (AIOps, intelligent alerting, automated diagnostics)
Hands-on experience with LLM-based developer/SRE productivity tools
Track record of driving AI adoption within engineering teams
Experience operating services at Apple-scale user volumes
Minimum Qualifications
10-15+ years of experience in SRE or adjacent disciplines (systems engineering, infrastructure engineering, production engineering), with at least 10 years in senior management roles
Demonstrated experience leading SRE organizations supporting large-scale, user-facing distributed services
Strong technical proficiency in Linux fundamentals, distributed systems concepts, networking, and infrastructure at scale
Demonstrated experience applying AI/ML tooling or LLM-based solutions to improve SRE or infrastructure operations
Ability to read and understand code produced by LLMs and evaluate its suitability for production use
Has defined or is actively executing an AI strategy for a large SRE organization
Proven ability to communicate at the executive level and negotiate across organizational boundaries
Pay & Benefits
At Apple, base pay is one part of our total compensation package and is determined within a range. This provides the opportunity to progress as you grow and develop within a role. The base pay range for this role is between $267,800 and $401,700, and your base pay will depend on your skills, qualifications, experience, and location.
Apple employees also have the opportunity to become an Apple shareholder through participation in Apple's discretionary employee stock programs. Apple employees are eligible for discretionary restricted stock unit awards, and can purchase Apple stock at a discount if voluntarily participating in Apple's Employee Stock Purchase Plan. You'll also receive benefits including: Comprehensive medical and dental coverage, retirement benefits, a range of discounted products and free services, and for formal education related to advancing your career at Apple, reimbursement for certain educational expenses - including tuition. Additionally, this role might be eligible for discretionary bonuses or commission payments as well as relocation. Learn more about Apple Benefits
Note: Apple benefit, compensation and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.

What Apple employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Apple logo

About Apple

Sourced by ZipRecruiter

Imagine what you could do here! At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish. Dynamic, intelligent people and inspiring, innovative technologies are the norm here. The people who work here have reinvented entire industries with all Apple Hardware products. The same real passion for innovation that goes into our products also applies to our practices strengthening our dedication to leave the world better than we found it.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Cupertino, CA, US

Year founded

1976