1

Ai Reliability Engineer Jobs in California (NOW HIRING)

About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...

About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...

About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...

Head of SRE

Palo Alto, CA · On-site

$67 - $89.25/hr

Wand AI is a company focused on integrating AI into the workforce, enabling humans and AI agents to work together efficiently. They are seeking a hands-on Head of SRE to establish and lead their Site ...

Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

About Runloop Runloop.ai is pioneering the next generation of infrastructure and orchestration to ... As a SRE, you'll be responsible for the reliability, observability, performance, and security of ...

We're a family-founded company on a mission to create the world's first AI-powered Personal ... The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the ...

Reliability Engineer

Costa Mesa, CA · On-site

$108K - $136K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

Reliability Engineer

Costa Mesa, CA · On-site

$110K - $138K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

As a Site Reliability Engineer, you will strengthen infrastructure, optimize tooling, deepen ... You will harness AI-assisted development and operational workflows to minimize toil, accelerate ...

next page

Showing results 1-20

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI Reliability Engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are AI Reliability Engineers?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges Ai Reliability Engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What job categories do people searching Ai Reliability Engineer jobs in California look for? The top searched job categories for Ai Reliability Engineer jobs in California are:
What cities in California are hiring for Ai Reliability Engineer jobs? Cities in California with the most Ai Reliability Engineer job openings:
Infographic showing various Ai Reliability Engineer job openings in California as of July 2026, with employment types broken down into 75% Full Time, 22% Part Time, and 3% Contract. Highlights an 71% Physical, 3% Hybrid, and 26% Remote job distribution.
Member of Technical Staff, AI Reliability & Monitoring Engineering Lead

Member of Technical Staff, AI Reliability & Monitoring Engineering Lead

Postman

San Francisco, CA

$256K - $276K/yr

Other

Posted 19 days ago


Job description

The Opportunity

Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman's AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally.

What You'll Do
  • Develop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features

  • Implement comprehensive observability and monitoring systems for real-time performance and fault detection

  • Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure

  • Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation

  • Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals

  • Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence

  • Drive continuous improvement in deployment practices, monitoring approaches, and incident management processes

About You
  • Have a strong background in AI reliability engineering, SRE, or DevOps for distributed systems

  • Understand the unique challenges of maintaining large-scale AI systems and integrating AI-specific metrics into reliability frameworks

  • Are experienced with cloud platforms, monitoring tools, and incident response automation

  • Are comfortable collaborating across teams to influence best practices for AI system reliability and operational health

  • Thrive in dynamic, fast-paced environments focusing on delivering reliable, safe AI-powered services

Bonus Skills and Experiences

  • Hands-on experience with AI/ML infrastructure, including GPU/xPU optimization and scaling

  • Familiarity with API platform operations and large-scale distributed services

  • Prior experience building or operating observability tools tailored for AI and agentic systems

  • Contribution to open-source projects or reliability engineering thought leadership

The reasonably estimated base salary for this role ranges from $256,000 to $276,000, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience.