1

Ai Reliability Engineer Jobs in California (NOW HIRING)

Reliability Engineer

Costa Mesa, CA · On-site

$108K - $136K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

Reliability Engineer

Costa Mesa, CA · On-site

$110K - $138K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

About Runloop Runloop.ai is pioneering the next generation of infrastructure and orchestration to ... As a SRE, you'll be responsible for the reliability, observability, performance, and security of ...

Site Reliability Engineer (SRE)

San Francisco, CA · On-site

$67.25 - $89.25/hr

Air Apps is a family-founded company on a mission to create the world's first AI-powered Personal & Entrepreneurial Resource Planner. As a Site Reliability Engineer (SRE), you will be responsible for ...

Site Reliability Engineer (SRE)

Palo Alto, CA · On-site

$67 - $89.25/hr

Job Summary : Mithril is an AI infrastructure platform focused on making GPU compute more accessible and affordable. The Site Reliability Engineer (SRE) will contribute to the stability and ...

Staff Reliability Engineer

Santa Clara, CA · On-site +1

$67.25 - $89.50/hr

Join us to put AI to work for people. Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and ...

Staff Reliability Engineer

Santa Clara, CA · On-site

$67.25 - $89.50/hr

Join us to put AI to work for people. Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and ...

Site Reliability Engineer (SRE)

San Francisco, CA · On-site

$67.25 - $89.25/hr

Job Summary : Mithril is an AI infrastructure platform focused on making GPU compute more accessible and affordable. The Site Reliability Engineer (SRE) will contribute to the stability and ...

Site Reliability Engineer

Cupertino, CA · On-site

$70.25 - $93.50/hr

VITURE is the #1 XR glasses brand in the US, aiming to create the first great AI interface you wear. The Site Reliability Engineer will design and maintain scalable cloud infrastructure for ...

You'll be building and operating the core systems that power agentic AI at scale. Your mission ... Partner with platform engineers to ensure reliability is designed into new features from day one.

Staff Site Reliability Engineer

San Francisco, CA · On-site +1

$67.25 - $89.25/hr

Develop AI-powered infrastructure automation, including Kubernetes lifecycle management, IaC ... Staff Site Reliability Engineer (IV) Senior Site Reliability Engineer (III) What you'll bring to ...

New

Develop AI-powered infrastructure automation, including Kubernetes lifecycle management, IaC ... Staff Site Reliability Engineer (IV) Senior Site Reliability Engineer (III) What you'll bring to ...

Develop AI-powered infrastructure automation, including Kubernetes lifecycle management, IaC ... Staff Site Reliability Engineer (IV) Senior Site Reliability Engineer (III) What you'll bring to ...

New

AI aids engineers by speeding context gathering, clarifying reasoning, and reducing repetition, but it does not replace accountability. * Collaborate with SRE, product engineering, infrastructure ...

Execute reliability qualification testing for our highly integrated photonics-based AI platform and ... Bachelor's degree in Electrical Engineering, Computer Engineering or a similar field * 5+ years of ...

New

Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Blaxel is a new kind of cloud computing infrastructure optimized for agentic AI, and they are seeking a world-class Site Reliability Engineer. The role involves ensuring the reliability, performance ...

Showing results 21-40

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What job categories do people searching Ai Reliability Engineer jobs in California look for? The top searched job categories for Ai Reliability Engineer jobs in California are:
What cities in California are hiring for Ai Reliability Engineer jobs? Cities in California with the most Ai Reliability Engineer job openings:
Infographic showing various Ai Reliability Engineer job openings in California as of August 2026, with employment types broken down into 67% Full Time, 28% Part Time, 2% Temporary, and 3% Contract. Highlights an 67% Physical, 3% Hybrid, and 30% Remote job distribution.

Reliability Engineer

Anduril Industries

Costa Mesa, CA • On-site

$108K - $136K/yr

Other

Re-posted 15 days ago


Anduril rating

9.1

Company rating: 9.1 out of 10

Anduril

Based on 17 frontline employees who took The Breakroom Quiz

7.4

Company rating compared to similar companies: 7.4 out of 10

Manufacturers average

Based on 75,035 frontline employees who took The Breakroom Quiz


Job description

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.
ABOUT THE TEAM
The Reliability Engineering team partners across Anduril's engineering, manufacturing, and operations organizations to ensure our autonomous systems survive real-world conditions and deliver consistent performance for our customers. Our mandate spans the complete product lifecycle - from early concept design through production scale-up and fielded operations. We use data, statistical methods, and engineering analysis to forecast how our systems will perform over time and influence design decisions, including part selection, redundancy strategies, and maintenance approaches. We establish tailored reliability frameworks by program and partner with Quality to drive design robustness, failure prevention, and rapid response to issues in the field.
ABOUT THE JOB
Anduril's Reliability & System Safety Engineering organization is hiring Reliability Engineers across our product portfolio. Domains may include Air Dominance & Strike, Counter-UAS, Maritime, Ground Systems, Space, and Mission Autonomy. These are highly collaborative, hands-on roles spanning the entire product lifecycle - from early concept trade studies and detailed design through manufacturing scale-up, production, and sustained field performance. You'll be the reliability champion throughout the design and development process, working closely with mechanical, electrical, and software engineering teams to ensure designs are robust. You'll also partner with manufacturing, quality, test operations, and field service teams to transition these products from engineering prototypes to production-ready systems. Your scope will span performing predictive reliability analysis and FMEAs, developing and executing qualification and production test plans, leading root-cause failure investigations from both the factory floor and the field, and closing the loop to drive continuous design and process improvements. We are looking for candidates who bring deep technical knowledge of hardware reliability and failure mechanisms, practical experience in both design and production environments, and the ability to drive cross-functional alignment without direct authority. These roles demand independent, data-driven engineers who can build reliability infrastructure from the ground up and take ownership of product performance from concept to deployment.
WHAT YOU'LL DO
  • Support and review system block diagrams, interface control documents, and schematics for compliance, redundancy, and accuracy against requirements and concept of operations.
  • Lead Failure Modes & Effects Analysis (FMEA) and Fault Tree Analysis (FTA) efforts, collaborating with cross-functional engineering teams to implement design improvements.
  • Perform predictive reliability analysis to calculate probability of loss of control, loss of asset, and MTTF/MTBF utilizing Anduril requirements, specifications, and relevant industry standards (i.e. MIL-HDBK-217Plus).
  • Develop and execute comprehensive Qualification and Validation Test plans (including environmental, life/limit, and accelerated testing) to ensure appropriate requirements coverage and de-risking for production readiness; partner with Test Engineering to fixture and instrument test campaigns.
  • Work with Quality and Manufacturing Engineering to develop, refine, and validate assembly, integration, and testing processes (including inspection/acceptance test gates) to optimize build reliability and reduce time-to-deploy.
  • Develop Maintainability prediction models using MIL-HDBK-472, and support the development of Preventative Maintenance & Sparing plans.
  • Perform Weibull analysis utilizing data from qualification campaigns, production testing, and fielded assets to predict failures and identify trends.
  • Lead failure investigations on anomalies originating from design validation, production testing, and field returns; drive root cause analysis and corrective/preventative actions (CAPA) to closure.
  • Support Development Milestone Reviews through the identification of appropriate entry and exit criteria, guiding the transition of designs from engineering development to high-rate production.
  • Develop and implement standardized reliability processes, tools, and infrastructure to streamline and accelerate our design, manufacturing, and deployment operations.
REQUIRED QUALIFICATIONS
  • Minimum of 5 years experience as a Reliability Engineer, Design/Development Engineer, or Test Engineer
  • B.S. Degree in Mechanical Engineering, Aerospace Engineering, Systems Engineering, or equivalent technical discipline
  • Experience with safety-critical hardware & software systems in the defense or aerospace industry
  • Experience or familiarity with MIL-STD-810, MIL-HDBK-217, MIL-STD-461, MIL-STD-516C, and MIL-STD-1629
  • Experience setting up a Reliability Engineering framework for a product or program, including support of production testing, field testing, data acquisition, and data analysis
  • Expertise in reliability analysis techniques, including FMEA/FMECA, predictive modeling, Weibull analysis, and Fault Tree Analysis
  • Experience with generating qualification testing requirements and executing tests to proactively characterize product behavior and performance
  • Failure investigation/analysis experience with proven track record of solving problems and preventing reoccurrence
  • Design Review Experience (PDR, CDR, MRR) and confidence in presentation skills across all levels of leadership
  • Eligibility to obtain/maintain a U.S. Secret clearance, as required
PREFERRED QUALIFICATIONS
  • M.S. Degree in Mechanical Engineering, Aerospace Engineering, Systems Engineering, or equivalent technical discipline
  • Experience with MIL-HDBK-217Plus, MIL-HDBK-472, MIL-STD-1916
  • Experience with Manufacturing Readiness Level (MRL) Assessments, Manufacturing Review Board (MRB), Quality Inspection, Tooling, or Calibration
  • Experience with risk management, change control/change management reviews, and software/firmware HITL/SITL
  • Experience with HALT (Highly Accelerated Life Testing), HASS (Highly Accelerated Stress Screening), Environmental Testing, and/or Mechanical Testing
  • Technical writing experience developing requirements, standards, specifications, user guides, and policies

US Salary Range
$146,000-$194,000 USD
The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including:
Benefits
At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you're supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.
Protecting Yourself from Recruitment Scams
Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information.
To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind:
  • No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates.
  • Please always verify communications:
    • Direct from Anduril: If you receive an email from one of our recruiters, it will only come from an @anduril.com address.
    • Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to
  • Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain is @anduril.com before providing any personal information or clicking on links.
  • What to Do If You Suspect Fraud: Should you encounter any questionable or fraudulent outreach claiming to be from Anduril, please report it immediately to Your proactive caution is invaluable in protecting your personal information and upholding the security and trustworthiness of our recruitment efforts.

Data Privacy
To view Anduril's candidate data privacy policy, please visit
By submitting your application, you consent to Anduril Industries using a third-party service provider to conduct pre-employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third-party service provider provides risk-intelligence services that may include analysis of sanctions and watchlists, adverse media, public-record information, and other lawful open-source or commercial data sources. This third-party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.

Working at Anduril

About Anduril, in their own words

From Anduril

As the world enters an era of strategic competition, Anduril delivers cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the warfighter in months, not years.


What Anduril employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Anduril Industries logo

About Anduril Industries

Sourced by ZipRecruiter

Anduril Industries is a trailblazer in the technology industry based in Costa Mesa, CA, US. Founded in 2017 by Palmer Luckey, the creator of Oculus VR, the company focuses on developing innovative technology to equip and empower those in the defense sector. Its primary products include cutting-edge autonomous systems and AI software that assist in combating threats to national and global security. The mission of Anduril Industries is to integrate technology and defense by building transformative, scalable solutions that ensure a safer world.

Industry

Guided missile and space vehicle manufacturing

Company size

501 - 1,000 Employees

Headquarters location

Costa Mesa, CA, US

Year founded

2017

Social media