1

Ai Reliability Engineer Jobs in California (NOW HIRING)

Join us to help build the next generation of AI hardware solutions. You will be part of a highly ... With a startup-like culture, we move quickly and give engineers the opportunity to drive ...

New

Site Reliability Engineer (SRE)

Palo Alto, CA ยท On-site

$67 - $89.25/hr

Mithril is an AI infrastructure platform focused on making GPU compute accessible for leading enterprises and AI startups. The Site Reliability Engineer will contribute to the stability and ...

Site Reliability Engineer

Newport Beach, CA ยท On-site

$61.25 - $81.50/hr

They are seeking a Site Reliability Engineer to support and maintain the service quality of their ... Obsidian Security provides SaaS and AI security software that detects threats, manages risks, and ...

Site Reliability Engineer

Palo Alto, CA ยท On-site

$67 - $89.25/hr

They are seeking a Site Reliability Engineer to support and maintain the service quality of their ... Obsidian Security provides SaaS and AI security software that detects threats, manages risks, and ...

Join NVIDIA, an innovator in computer graphics, PC gaming, and accelerated computing, as we step into the next era shaped by AI. As a Senior Reliability Engineer, you'll work within a focused team ...

Staff Reliability Engineer

Santa Clara, CA ยท On-site

$67.25 - $89.50/hr

The combination brings together Veza's AI-native Access Graph with ServiceNow's AI Control Tower ... We are seeking an exceptional Staff Site Reliability Engineer to lead critical infrastructure ...

Site Reliability Engineer (SRE)

Palo Alto, CA ยท On-site

$67 - $89/hr

About Mithril Mithril is an AI infrastructure platform built to make GPU compute more accessible ... What Makes This Role Different Most SRE roles at this stage are primarily reactive - on-call, ...

Site Reliability Engineer

Santa Clara, CA ยท On-site

$230K - $250K/yr

It's the foundation for autonomous networking, giving engineers and AI agents the ability to know ... Forward is looking for a Site Reliability Engineer About the Role This is not a "keep the lights on ...

Hardware Reliability Engineer

Mountain View, CA ยท On-site

$120K - $152K/yr

Own the reliability of the advanced packages and systems that turn our AI accelerator silicon into products that survive years in the field. You'll define how we qualify 2.5D/3D and heterogeneously ...

Senior Site Reliability Engineer

San Francisco, CA ยท On-site

$67.25 - $89.25/hr

Programming in Python supported by Gen AI tooling to accelerate development of mission critical ... Site Reliability Engineering Performance Indicators The main KPIs that aid in understanding the ...

Showing results 41-60

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What job categories do people searching Ai Reliability Engineer jobs in California look for? The top searched job categories for Ai Reliability Engineer jobs in California are:
What cities in California are hiring for Ai Reliability Engineer jobs? Cities in California with the most Ai Reliability Engineer job openings:
Infographic showing various Ai Reliability Engineer job openings in California as of August 2026, with employment types broken down into 67% Full Time, 28% Part Time, 2% Temporary, and 3% Contract. Highlights an 67% Physical, 3% Hybrid, and 30% Remote job distribution.

Staff Software Engineer - Reliability (US Citizen Only)

Rubrik Job Board

Palo Alto, CA โ€ข On-site

$67 - $89/hr

Other

Re-posted 9 days ago


Job description

About Team & About Role

The Site Reliability Engineering (SRE) team at Rubrik ensures the absolute reliability, availability, performance, and security of our enterprise infrastructure services, spanning both global SaaS platforms and government-compliant environments. We operate at the intersection of software development and systems engineering, prioritizing hyperscale platform automation, self-healing architectures, and structural resiliency. As a Staff Site Reliability Engineer, you will serve as a primary technical leader and architect across our broader distributed cloud systems. You will drive long-term technical roadmaps, establish cross-organizational reliability standards, and solve complex distributed systems challenges that safeguard both enterprise and public sector environments.ย 

Beyond the core SRE charter, this Staff role also leads the Application-SRE team - a US-based group that partners closely with engineering, Sales, and Support to unblock POCs, drive complex customer escalations to resolution, and convert recurring field signals into engineering and reliability roadmap items. You will be the technical leader and project owner for Application-SRE: setting direction, tracking commitments, and ensuring the team operates as a high-leverage bridge between the field and the broader engineering org.

What You'll Do

As a Staff Site Reliability Engineer, you will possess engineering-wide influence and take ownership of the following critical areas:

  • Infrastructure Strategy & Architecture: Formulate and execute the architectural vision for Rubrik's Cloud Platform, optimizing backend infrastructure systems like Kubernetes, MySQL, and cloud-native services for performance, security, and multi-region scale.
  • Hyperscale Automation & Platform Tooling: Build, scale, and maintain sophisticated custom internal tools, platform controllers, and automation frameworks in Go or Python to systematically eliminate operational toil.
  • AI Infrastructure for SaaS: Deploy, scale, and operate the AI infrastructure that powers Rubrik's SaaS offerings, owning the reliability, performance, cost, and security controls required to run AI workloads in multi-tenant, compliance-bound environments.
  • AI for SRE & Engineering Productivity: Drive the adoption of AI-driven solutions across the SRE charter to compress toil and multiply the org - applying agentic and LLM-based approaches to automated triage, incident response, operational analysis, and developer productivity.
  • AI Adoption Guardrails for SaaS Reliability: Build the guardrails, controls, and platform patterns that keep Rubrik's SaaS reliable as AI adoption accelerates across product and engineering, ensuring new AI capabilities ship without eroding availability, performance, security, or cost posture.
  • Cross-Functional Leadership: Wield engineering-wide influence to create technical consensus among component, platform, and security engineering teams, effectively "shifting left" to embed structural resilience, capacity guards, and compliance from initial feature designs.
  • Reliability Governance: Define, audit, and enforce robust Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets across all critical enterprise platform services, translating telemetry insights into actionable product roadmaps during executive reviews.
  • Incident Command & Operations Review: Serve as a primary Incident Commander for high-severity cloud outages, establishing roles, directing mitigation vectors under pressure, and orchestrating comprehensive, blameless post-mortems that drive durable systemic fixes.
  • Cost Governance & Capacity Modeling: Architect cost-observability tools and attribution frameworks, leading cloud infrastructure capacity forecasting, resource quota optimization, and vendor SLA management.
  • Application-SRE Leadership: Set the technical direction for the Application-SRE team, raising the bar on how the team diagnoses, mitigates, and durably resolves the most complex customer-impacting issues across our platform.
  • Technical Multiplier & Mentorship: Champion SRE best practices, mentoring senior and junior individual contributors across the organization, participating in interview frameworks, and actively raising the collective technical bar.
  • On-Call Rotations: Participate in on-call rotations
Experience You'll Need
  • Citizenship & Residency: Must be a US Citizen currently residing on CONUS soil (strict regulatory requirement to enable support for federal and FedRAMP environments when required).
  • Education: BS, MS, or PhD in Computer Science, Computer Engineering, or a highly related technical discipline.
  • Industry Experience: A minimum of 8-12+ years of software engineering and production cloud infrastructure experience, with at least 5+ years dedicated to a formal SRE, DevOps, or Platform engineering role operating hyperscale SaaS products.
  • Technical Depth: Comprehensive, hands-on programming expertise in Golang, Python, or Java with a deep grasp of concurrency models, data structures, and test-driven software design patterns.
  • Distributed Systems Expertise: Proven proficiency designing, deploying, analyzing, and auditing complex, large-scale distributed systems, database topologies, and high-availability public cloud meshes.
  • Systems Internals: Authoritative operational command of Unix/Linux operating system environments (process models, file systems, kernels), systems administration, and advanced L4/L7 networking protocols.
  • AI Systems Fluency: Working knowledge of operating AI systems in production - including model serving, cost trade-offs, and the reliability and safety considerations of LLM- and agent-based workloads. Practical judgment on when AI is the right tool versus deterministic automation.
  • Field-to-Product Feedback Loop: Institutionalize the channel that converts patterns from customer escalations and POCs into prioritized product and reliability feedback, partnering directly with Product, Sales Engineering and Support leadership.
  • Customer & Field Fluency: Track record of partnering directly with Sales, Support, and customers on escalations and POCs, and translating field signals into engineering action.
  • Leadership Capability: Demonstrated history of technical leadership, mapping architectural dependencies, managing multi-team technical projects, and guiding organizations through critical platform shifts with high technical judgment.
Preferred Qualifications
  • Extensive production experience provisioning, lifecycle-managing, and recovering enterprise-scale Kubernetes (GKE, EKS) deployments and large-scale relational/non-relational databases (MySQL).
  • Prior experience building, certifying, or auditing infrastructure environments under compliance structures such as FedRAMP (High/Moderate), SOC 2, ISO 27001, or CJIS.
  • Fluency in Infrastructure-as-Code (Terraform, Pulumi) module design, multi-tenant state isolation, and enterprise observability fabrics (Prometheus, Grafana, OpenTelemetry).
  • Exposure to building AI- or LLM-powered internal tooling and applying it to SRE, operations, or engineering productivity use cases.
  • Familiarity with the operational considerations of running AI workloads on cloud and Kubernetes platforms.