1

Observability Jobs in California (NOW HIRING)

The Observability Engineering organization at CoreWeave is responsible for the platforms and practices that help engineers understand, operate, and improve production systems at scale. This team owns ...

Showing results 21-40

Observability information

See California salary details

$16

$59

$85

How much do observability jobs pay per hour?

As of Aug 22, 2026, the average hourly pay for observability in California is $59.74, according to ZipRecruiter salary data. Most workers in this role earn between $50.05 and $68.56 per hour, depending on experience, location, and employer.

What is an observability?

An Observability job focuses on ensuring the performance, reliability, and health of software systems by collecting, analyzing, and visualizing telemetry data such as logs, metrics, and traces. Professionals in this field work with monitoring tools, distributed tracing, and alerting systems to detect and troubleshoot issues proactively. They collaborate with engineering and operations teams to improve system visibility, reduce downtime, and enhance overall system performance.

What does an observability do?

In an Observability role, your daily tasks often include designing and maintaining monitoring dashboards, configuring alerts, analyzing system logs, and working closely with development and operations teams to troubleshoot issues. You'll proactively identify areas of improvement to increase system reliability, document monitoring strategies, and support incident response efforts. Collaboration is key, as you may participate in post-incident reviews and help drive architectural improvements based on the data you collect. The role is dynamic and requires a proactive approach to ensure systems stay healthy and downtime is minimized.

What are the key skills and qualifications needed to thrive in an observability role?

To thrive in an Observability role, you need a strong background in monitoring, alerting, logging, and analyzing system performance, often supported by a degree in computer science or related field. Familiarity with tools such as Prometheus, Grafana, Datadog, Splunk, and experience with cloud platforms and scripting languages is crucial. Excellent problem-solving, communication, and collaboration skills help you work effectively with cross-functional engineering and operations teams. These capabilities are essential to ensure system reliability, quickly detect issues, and maintain seamless digital experiences.

Is observability a good career?

Observability is a growing field within IT and software engineering that involves monitoring, logging, and analyzing system performance using tools like Prometheus, Grafana, and Elasticsearch. It offers opportunities for specialization, high demand for skills, and roles in DevOps and site reliability engineering, making it a viable career choice for those interested in system reliability and automation.

What are the most commonly searched types of Observability jobs in California?

The most popular types of Observability jobs in California are:

What job categories do people searching Observability jobs in California look for?

The top searched job categories for Observability jobs in California are:

What cities in California are hiring for Observability jobs?

Cities in California with the most Observability job openings:

Infographic showing various Observability job openings in California as of August 2026, with employment types broken down into 97% Full Time, and 3% Contract. Highlights an 73% Physical, 8% Hybrid, and 19% Remote job distribution, with an average salary of $124,260 per year, or $59.7 per hour.

Principal Observability Platform Engineer

Recruitment.ai

San Francisco, CA • On-site

Other

Posted 9 days ago


Job description

ABOUT THE ROLE

We''re hiring a Principal/Staff Observability Platform Engineer to own the technical direction of our observability platform — the systems that give us deep visibility into GPU clusters, AI workloads, and the infrastructure running them.

This is a "define, build, and lead" role, not a "maintain and operate" role. You''ll set the architectural roadmap, raise the engineering bar across teams, and make sure the platform scales ahead of the business, not behind it. We have a strong bias toward simplicity — the systems you build should be easy to operate and self-evidently correct when something goes wrong.

RESPONSIBILITIES

- Own the technical strategy and architecture for observability across metrics, logs, traces, and alerting at scale
- Drive platform decisions with multi-year impact: tooling, data models, ingestion patterns, retention, cardinality management
- Identify systemic gaps before they become incidents; design platforms that make failure visible and fast to diagnose
- Partner with SRE, infrastructure, and AI/ML teams to embed observability natively into how we build and operate
- Define standards and patterns that other engineers adopt because they''re clearly better, not by mandate
- Mentor and technically grow the observability team
- Lead incident postmortems and drive durable platform improvements
- Evaluate and introduce tooling that improves signal quality, operational efficiency, or scalability — and retire what doesn''t

REQUIRED SKILLS & EXPERIENCE

- 8+ years in SRE, infrastructure engineering, platform engineering, or observability-focused roles
- Proven experience operating observability infrastructure at serious scale
- Deep hands-on experience with a significant subset of: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic
- Strong engineering fundamentals in Python, Go, or similar
- Kubernetes at scale
- Infrastructure-as-Code as default practice (Terraform, Ansible, or equivalent)
- Demonstrated ability to architect systems, write code, review others'' work, and clearly explain tradeoffs
- Track record of influencing engineering direction across teams without formal authority

PREFERRED SKILLS & EXPERIENCE

- Experience with high-volume streaming pipelines for observability data (Kafka, Vector, Fluent Bit, etc.)
- Background in AI/ML infrastructure observability: GPU utilization, training job visibility, inference latency
- Familiarity with GPU infrastructure or HPC environments (Slurm)
- Prior experience defining observability strategy at an organizational level

EQUAL OPPORTUNITY

We strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent individuals, parents, carers, and people from lower socio-economic backgrounds. If there''s anything we can do to accommodate your specific situation, please let us know.