1

Observability Aiops Engineer Jobs in California (NOW HIRING)

Senior AIOps ML Engineer

Los Angeles, CA · On-site

$112K - $154K/yr

The engineer will also focus on security and compliance observability, collaborating with security ... ML Model Development & AIOps: Design, train, and deploy machine learning models for streaming ...

Senior Site Reliability Engineer, AIOPs

Santa Clara, CA · On-site

$67 - $89/hr

Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing ... Proven programming experience building automation tools or services - ideally in Python, or similar ...

Senior AIOps ML Engineer

Woodland Hills, CA · On-site

$110K - $151K/yr

The role involves designing and developing machine learning models and data engineering solutions for AIOps, focusing on multi-domain observability data and real-time anomaly detection.

Engineer automation workflows to reduce manual effort and improve signal quality. * Collaborate ... Strong background integrating monitoring tools into a central observability/AIOps platform. * Hands ...

Senior Sales Engineer

Pleasanton, CA · On-site

$300K - $330K/yr

A background in the following domains: vAI, Observability, AIOps, automation, and enterprise operations. * Master's degree in Engineering preferred. * Consistent history of achieving or surpassing ...

Senior Engineer - ML

Pleasanton, CA · On-site +1

$116K - $159K/yr

... AIOps platform within the Observability product. The candidate will work closely with the Lead ... Collaborate with platform and engineering teams to build scalable model-serving and agent-serving ...

New

Senior Engineer - ML

Pleasanton, CA · On-site +1

$115K - $158K/yr

Experience in AIOps, Observability, SRE, IT operations, or incident management domains ... Experience applying AI/ML to RCA, anomaly explanation, incident summarization, service health ...

next page

Showing results 1-20

Observability Aiops Engineer information

What are some common challenges faced by Observability AIOps Engineers in integrating monitoring solutions across diverse technology stacks?

Observability AIOps Engineers often encounter challenges when integrating monitoring and analytics tools across a mix of legacy systems, cloud-native applications, and various third-party platforms. Ensuring consistent data collection, normalization, and visualization can be complex due to differing protocols, data formats, and tool compatibility. Collaboration with development, operations, and security teams is crucial to address these challenges, streamline workflows, and maintain a unified observability platform. Staying current with evolving AIOps technologies and best practices is also vital for continued success in this dynamic role.

What is an Observability Aiops Engineer?

An Observability Aiops Engineer is a technology professional who focuses on implementing and managing observability tools and practices, often leveraging artificial intelligence for IT operations (AIOps). Their role is to ensure system reliability, performance, and uptime by monitoring, analyzing, and automating responses to IT incidents. They integrate data from logs, metrics, and traces to gain real-time insights, helping organizations quickly detect and resolve issues. This role combines expertise in software engineering, monitoring solutions, automation, and machine learning to improve the overall health and efficiency of IT environments.

What are the key skills and qualifications needed to thrive as an Observability AIOps Engineer, and why are they important?

To thrive as an Observability AIOps Engineer, you need expertise in systems monitoring, data analytics, automation, and a strong understanding of IT infrastructure, often supported by a degree in computer science or a related field. Familiarity with tools like Prometheus, Grafana, ELK stack, Splunk, and AIOps platforms, as well as certifications in cloud solutions (AWS, Azure, or GCP), are typically required. Strong problem-solving skills, collaboration, and a proactive mindset help you stand out in identifying and addressing system anomalies. These skills and qualities are crucial for maintaining high system reliability, reducing downtime, and enabling data-driven decision-making in complex IT environments.

What is the difference between Observability Aiops Engineer vs Site Reliability Engineer?

AspectObservability Aiops EngineerSite Reliability Engineer
Primary FocusMonitoring, analyzing, and improving system observability using AI and automationEnsuring system reliability, scalability, and performance of services
Skills & CertificationsKnowledge of AI/ML, monitoring tools, scripting, cloud platformsSystems engineering, scripting, cloud infrastructure, incident management
Work EnvironmentDevOps teams, monitoring platforms, AI toolsOperations, development teams, cloud environments
Industry UsageTech companies, cloud providers, organizations focusing on AI-driven monitoringLarge-scale tech firms, SaaS providers, internet services

While both roles focus on system performance and reliability, the Observability Aiops Engineer specializes in leveraging AI and automation to enhance system observability, whereas the Site Reliability Engineer concentrates on maintaining overall system stability and scalability. Both roles often collaborate but have distinct core responsibilities.

What are popular job titles related to Observability Aiops Engineer jobs in California? For Observability Aiops Engineer jobs in California, the most frequently searched job titles are:
What job categories do people searching Observability Aiops Engineer jobs in California look for? The top searched job categories for Observability Aiops Engineer jobs in California are:
What cities in California are hiring for Observability Aiops Engineer jobs? Cities in California with the most Observability Aiops Engineer job openings:
Infographic showing various Observability Aiops Engineer job openings in California as of July 2026, with employment types broken down into 67% Full Time, and 33% Contract. Highlights an 100% In-person job distribution.

Senior Site Reliability Engineer, AIOPs

NVIDIA

Santa Clara, CA • On-site

$67 - $89/hr

Full-time

Re-posted 5 days ago


Nvidia rating

9.6

Company rating: 9.6 out of 10

Based on 17 frontline employees who took The Breakroom Quiz

8th of 241 rated software companies


Job description

Job Summary:
NVIDIA has been transforming computer graphics and computing for over 25 years, and they are seeking a Senior Site Reliability Engineer to join their innovative team. The role involves operating an AI Data Center AIOps platform, ensuring uptime, performance, and data integrity while collaborating with engineering teams to create actionable insights and automation.
Responsibilities:
• Continuously monitor platform health via dashboards/logs/metrics, automate recurring checks, and keep reliability + resource efficiency on track.
• Own Kubernetes deployments end-to-end (runbooks, canary checks, post-deploy validation), and lead rollbacks/remediations when needed.
• Lead first-level incident triage: collect diagnostics, identify likely root causes, and hand off clear, actionable findings to engineering.
• Build and maintain runbooks/SOPs/checklists, pushing continuous improvement through automation.
• Manage deployment infrastructure and packaging (Helm + Terraform/IaC) to keep environments scalable, consistent, and reproducible.
• Contribute in adjacent functional areas to grow and help your team members!
Qualifications:
Required:
• BS/MS in CS/CE (or equivalent experience) and 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.
• Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing incidents, and follow-up evaluations that drive measurable improvements.
• Deep Kubernetes + containers experience (deploying, debugging, scaling) for telemetry-heavy microservices—ingestion, processing, storage, APIs, and UI.
• Automation-first approach: solid scripting (Python/Bash), CI/CD, and infrastructure-as-code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil.
• Clear communicator who writes excellent runbooks/docs and can translate ambiguous requirements into concrete operational practices and dependable customer-facing reliability.
Preferred:
• Strong Linux + networking fundamentals, distributed systems instincts, and hands-on ops for Kubernetes/services/streaming stacks are ideal; bonus for experience with observability platforms at scale.
• Experience building safe automation that operators trust: canary releases, automated rollback criteria, 'monitoring for the monitoring' (lag/drop/error budgets), and replay/backfill pipelines with correctness checks.
• Strong in distributed/streaming systems operations (Kafka/Pulsar, Flink/Spark, ClickHouse/Elastic/TSDBs, object storage)—and can reason about backpressure, hotspots, and failure domains end-to-end.
• Proven programming experience building automation tools or services — ideally in Python, or similar languages — to simplify operations and scale recurring processes.
• Proven experience running large‑scale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with hands‑on experience with observability tools — you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar.
Company:
NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI. Founded in 1993, the company is headquartered in Santa Clara, USA, with a team of 10001+ employees. The company is currently Late Stage.

What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US

Year founded

1993