1

Contractual Observability Engineer Jobs in California

Site Reliability Engineer

San Francisco, CA ยท Remote

$67.25 - $89.25/hr

The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution. Key responsibilities * Define and operate Service Level ...

Sr Staff Engineer, Data Infrastructure

San Jose, CA ยท On-site

$127K - $172K/yr

... environments. โ€ข Observability at Scale: Build high-cardinality monitoring systems using ... contractual, and/or regulatory requirements. Company : Archer is an aerospace company that ...

Sr Staff Site Reliability Engineer

San Jose, CA ยท On-site

$207K - $259K/yr

Develop and implement comprehensive observability strategies, including monitoring, logging, and ... ITAR, contractual, and/or regulatory requirements. Please note that this is intended to provide a ...

Sr Staff Site Reliability Engineer

San Jose, CA ยท On-site

$207K - $259K/yr

Develop and implement comprehensive observability strategies, including monitoring, logging, and ... ITAR, contractual, and/or regulatory requirements. Please note that this is intended to provide a ...

Engineer VI

Poway, CA ยท On-site

$116K - $208K/yr

... contractual obligations. The engineer will also stay informed of industry trends and advance ... low observability (LO) features. * Collaborate with the external aerodynamics team to improve ...

Observability at Scale: Build high-cardinality monitoring systems using VictoriaMetrics and Vector ... ITAR, contractual, and/or regulatory requirements. Please note that this is intended to provide a ...

Partner with security and engineering leadership to define priorities, investment areas, and ... On security terms, third-party risk, and contractual obligations. * Executive Leadership: Senior ...

You will work closely with engineering, infrastructure, security, customer support, and partner ... Continuously improve monitoring, alerting, and observability tooling to reduce noise and increase ...

$200K - $280K/yr

Leveraging our holistic portfolio of capabilities in consulting, design, engineering, and ... Emerging areas such as AI Data Centres and Observability * Translate client business challenges ...

$190K - $250K/yr

Leveraging our holistic portfolio of capabilities in consulting, design, engineering, and ... Emerging areas such as AI Data Centres and Observability * Translate client business challenges ...

Contractual Observability Engineer information

What is the difference between Contractual Observability Engineer vs Site Reliability Engineer?

AspectContractual Observability EngineerSite Reliability Engineer
Primary FocusImplementing observability tools, monitoring, and logging systemsEnsuring system reliability, scalability, and performance
Skills & CertificationsMonitoring tools, scripting, cloud platformsSystems engineering, automation, incident management
Work EnvironmentDevOps teams, cloud environments, monitoring platformsOperations teams, production systems, automation tools
Industry UsageTech companies, SaaS providers, cloud servicesLarge-scale tech firms, internet services, cloud providers

While both roles involve technical expertise in cloud and systems, a Contractual Observability Engineer primarily focuses on implementing and maintaining observability tools, whereas a Site Reliability Engineer emphasizes system reliability and performance. Understanding these differences helps organizations assign the right responsibilities and skills to each role.

What are the most commonly searched types of Observability Engineer jobs in California? The most popular types of Observability Engineer jobs in California are:
What are popular job titles related to Contractual Observability Engineer jobs in California? For Contractual Observability Engineer jobs in California, the most frequently searched job titles are:
What job categories do people searching Contractual Observability Engineer jobs in California look for? The top searched job categories for Contractual Observability Engineer jobs in California are:
What cities in California are hiring for Contractual Observability Engineer jobs? Cities in California with the most Contractual Observability Engineer job openings:

Site Reliability Engineer

STN Inc

San Francisco, CA โ€ข Remote

$67.25 - $89.25/hr

Full-time

Re-posted 17 days ago


Job description

Site Reliability Engineer

Platform and software ยท shared across customers

Reports to: Director, Site Reliability

Location: Remote (US)

Department: Cloud Platform Engineering / SRE/Reliability

Position summary

The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.

Key responsibilities
  • Define and operate Service Level Objectives (SLOs) aligned with customer SLAs

  • Build and maintain the observability stack including metrics, logs, traces, and alerting

  • Lead incident response and chair post-incident reviews

  • Drive automation to reduce toil and improve mean-time-to-recover (MTTR)

  • Author and maintain operational runbooks alongside the NOC

  • Manage on-call rotation, escalation paths, and incident-management tooling

  • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering

  • Drive chaos engineering, game days, and reliability testing programs

  • Produce SLA performance reports in coordination with the SLA Manager

  • Mentor junior engineers and contribute to engineering culture

Required qualifications
  • 5+ years in SRE, DevOps, or production engineering roles

  • Strong programming skills in Go, Python, or both

  • Hands-on experience operating Kubernetes-based platforms at scale

  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)

  • Strong incident management experience including major-incident command

Preferred qualifications
  • GPU or HPC platform operational experience

  • Familiarity with SLA-driven customer environments and credit calculations

  • Experience with chaos engineering tools (Gremlin, Litmus, or similar)

  • Published SRE content or contributions