1

Ai Reliability Engineer Jobs in California (NOW HIRING)

AI Inference Engineer

San Jose, CA ยท On-site

$134K - $161K/yr

This role is pivotal in advancing F5's AI capabilities, ensuring enterprise-grade reliability by ... Programming Languages: Proficiency in programming languages such as Python , C++ , Rust , or Golang ...

About the Role We are building a high-performance SRE function to support one of the world's fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help ...

Reliability Engineer

Sunnyvale, CA ยท On-site

$120K - $151K/yr

Analyzes preliminary plans and develops reliability engineering programs to achieve company ... in AI prompting as it applies to this position and practice Active or Current security clearance ...

Site Reliability Engineer

San Francisco, CA ยท On-site

$67.25 - $89.25/hr

About Runloop Runloop.ai is pioneering the next generation of infrastructure and orchestration to ... As a SRE, you'll be responsible for the reliability, observability, performance, and security of ...

Principal SRE - AI Inference

Sunnyvale, CA ยท On-site

$67 - $89/hr

About the Role We are building a high-performance SRE function to support one of the world's fastest-growing AI inference services, powered by the Wafer-Scale Engine (WSE). This team will help ...

Site Reliability Engineer (SRE)

Palo Alto, CA ยท On-site

$67 - $89.25/hr

Job Summary : Mithril is an AI infrastructure platform focused on making GPU compute more accessible and affordable. The Site Reliability Engineer (SRE) will contribute to the stability and ...

Site Reliability Engineer (SRE)

San Francisco, CA ยท On-site

$67.25 - $89.25/hr

Job Summary : Mithril is an AI infrastructure platform focused on making GPU compute more accessible and affordable. The Site Reliability Engineer (SRE) will contribute to the stability and ...

Site Reliability Engineer

Santa Clara, CA ยท On-site

$67.25 - $89.50/hr

Are you interested in working with the World's leading AI-first Quality Engineering Company? Ready ... We are looking for a Site Reliability Engineer to join our growing team in Riverwoods, IL United ...

Staff Reliability Engineer

Santa Clara, CA ยท On-site

$67.25 - $89.50/hr

Join us to put AI to work for people. Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and ...

Site Reliability Engineer

San Francisco, CA ยท On-site

$150 - $200/hr

About Runloop Runloop.ai is pioneering the next generation of infrastructure and orchestration to ... As a SRE, you'll be responsible for the reliability, observability, performance, and security of ...

Site Reliability Engineer

San Francisco, CA ยท On-site

$175K - $250K/yr

You'll be building and operating the core systems that power agentic AI at scale. Your mission ... Partner with platform engineers to ensure reliability is designed into new features from day one.

Site Reliability Engineer

San Francisco, CA

$67.25 - $89.25/hr

You'll be building and operating the core systems that power agentic AI at scale. Your mission ... Partner with platform engineers to ensure reliability is designed into new features from day one.

Reliability Engineer

Los Angeles, CA ยท On-site

$110K - $139K/yr

In the process, we are restoring America's ability to produce nuclear fuel to power AI, advanced ... About This Role As a Reliability Engineer at General Matter, you'll own the reliability of the ...

Reliability Engineer

Santa Clara, CA ยท On-site

$130K - $155K/yr

The primary purpose of the Reliability Engineer (focused on L10/L11 execution) is to own the physical, environmental, and mechanical stress-testing validation of integrated AI rack systems. This ...

Reliability Engineer

Los Angeles, CA ยท On-site

$110K - $139K/yr

In the process, we are restoring America's ability to produce nuclear fuel to power AI, advanced ... About This Role As a Reliability Engineer at General Matter, you'll own the reliability of the ...

Site Reliability Engineer

Cupertino, CA ยท On-site

$70.25 - $93.50/hr

VITURE is the #1 XR glasses brand in the US, aiming to create the first great AI interface you wear. The Site Reliability Engineer will design and maintain scalable cloud infrastructure for ...

Staff Site Reliability Engineer

San Francisco, CA ยท On-site +1

$67.25 - $89.25/hr

Develop AI-powered infrastructure automation, including Kubernetes lifecycle management, IaC ... Staff Site Reliability Engineer (IV) Senior Site Reliability Engineer (III) What you'll bring to ...

Meta is seeking a Reliability Engineer to drive the quality, durability, and long-term reliability ... Demonstrated use of AI tools to redesign reliability data analysis workflows, accelerate failure ...

Staff Site Reliability Engineer

San Francisco, CA ยท On-site +1

$67.25 - $89.25/hr

Develop AI-powered infrastructure automation, including Kubernetes lifecycle management, IaC ... Staff Site Reliability Engineer (IV) Senior Site Reliability Engineer (III) What you'll bring to ...

Showing results 41-60

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What job categories do people searching Ai Reliability Engineer jobs in California look for?

The top searched job categories for Ai Reliability Engineer jobs in California are:

What cities in California are hiring for Ai Reliability Engineer jobs?

Cities in California with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in California as of August 2026, with employment types broken down into 100% Full Time. Highlights an 60% In-person, and 40% Remote job distribution.

AI Inference Engineer

F5, Inc.

San Jose, CA โ€ข On-site

$134K - $161K/yr

Full-time

Posted 12 days ago


Job description

At F5, we strive to bring a better digital world to life. Our teams empower organizations across the globe to create, secure, and run applications that enhance how we experience our evolving digital world. We are passionate about cybersecurity, from protecting consumers from fraud to enabling companies to focus on innovation.
Everything we do centers around people. That means we obsess over how to make the lives of our customers, and their customers, better. And it means we prioritize a diverse F5 community where each individual can thrive.
Job Description
The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments-from GPU-rich data centers to resource-constrained edge devices-with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy.
This role is pivotal in advancing F5's AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance.
Key Responsibilities
High-Performance AI Serving
  • Build and maintain robust inference engines using tools like vLLM, TGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale.
  • Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications.

Hardware Acceleration and Optimization
  • Profile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs.
  • Collaborate with hardware teams to maximize utilization and performance across various computational environments.

Inference Orchestration and Scalability
  • Design and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration.
  • Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability.

Performance Monitoring and Observability
  • Establish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs).
  • Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale.

Technical Requirements
Required Skills:
  • Programming Languages: Proficiency in programming languages such as Python, C++, Rust, or Golang specifically for high-performance AI workflows.
  • Inference Tools: Proven hands-on experience with tools like vLLM, TensorRT, Llama.cpp, and Ollama for inference development and optimization.
  • Infrastructure Expertise: Strong familiarity with infrastructure technologies, including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure.
  • Hardware Optimization Expertise: Comprehensive understanding of GPU and AI hardware, including techniques for profiling and optimizing performance for accelerators like NVIDIA GPUs and TPUs.

Preferred Experience:
  • Prior experience deploying Large Language Models (LLMs) with advanced techniques like Speculative Decoding or PagedAttention.
  • Contributions to open-source inference libraries or hardware-level kernel development (e.g., CUDA, Triton kernels).
  • Background in MLOps or SRE roles focused on high-performance AI endpoints and reliability during demand surges.
  • Proficiency in designing scalable solutions for high-throughput inference environments optimized for traffic bursts.

Success Metrics (KPIs):
  • Latency Reduction: Continuously improve inference latency metrics, ensuring minimal Time to First Token (TTFT) and maximum tokens per second.
  • Cost Efficiency: Achieve lower "Cost per 1K Tokens" through better resource utilization and hardware optimization.
  • Scalability: Maintain system stability and reliability during traffic spikes, ensuring performance consistency across environments.
  • Throughput Maximization: Deploy models optimized for peak hardware usage and maximized process throughput.

Why Join F5?
F5 empowers you to push boundaries in AI optimization and high-performance engineering. Joining our team means:
  • Collaborating with cutting-edge technologies and hardware solutions to support real-time AI applications.
  • Advancing your career in a fast-paced, multidisciplinary environment focused on innovation, scalability, and problem-solving.
  • Driving transformative projects that deliver real-time AI reliability to global customers while maintaining cost and efficiency standards.
  • Working on advanced MLOps solutions that seamlessly scale enterprise AI systems and shape the future of intelligent deployment.

What Success Looks Like:
As an AI Inference Engineer at F5, success is measured by your ability to:
  • Combine technical expertise and problem-solving skills to deliver low-latency, scalable, and high-performing AI prediction systems.
  • Collaborate efficiently across cross-functional teams, participating in knowledge sharing and system refinement.
  • Demonstrate initiative by driving optimizations across hardware, tools, and orchestration processes, balancing immediate solutions with long-term architectural goals.
  • Translate complex AI and inference workflows into practical solutions that align with F5's strategic objectives.

The base pay range per annum for this position is: $176,600 - $265,000
F5 maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, geographic locations, and market conditions, as well as to reflect F5's differing products, industries, and lines of business. The pay range referenced is as of the time of the job posting and is subject to change. You may also be offered incentive compensation, bonus, restricted stock units, and benefits. More details about F5's benefits can be found at the following link: https://www.f5.com/company/careers/benefits. F5 reserves the right to change or terminate any benefit plan without notice.
#LI-ZB1
The Job Description is intended to be a general representation of the responsibilities and requirements of the job. However, the description may not be all-inclusive, and responsibilities and requirements are subject to change.
Please note that F5 only contacts candidates through F5 email address (ending with @f5.com) or auto email notification from Workday (ending with f5.com or @myworkday.com)
Equal Employment Opportunity
It is the policy of F5 to provide equal employment opportunities to all employees and employment applicants without regard to unlawful considerations of race, religion, color, national origin, sex, sexual orientation, gender identity or expression, age, sensory, physical, or mental disability, marital status, veteran or military status, genetic information, or any other classification protected by applicable local, state, or federal laws. This policy applies to all aspects of employment, including, but not limited to, hiring, job assignment, compensation, promotion, benefits, training, discipline, and termination. F5 offers a variety of reasonable accommodations for candidates. Requesting an accommodation is completely voluntary. F5 will assess the need for accommodations in the application process separately from those that may be needed to perform the job. Request by contacting accommodations@f5.com.