1

Observability Site Reliability Engineer Jobs in Riverside, CA

Staff AI Platform & Reliability Engineer

Santa Ana, CA · On-site

$172 - $220/hr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Lead reliability engineering: failure handling, degradation strategy, capacity and incident ... Strengthen deployment safeguards, observability and cost attribution * Set technical standards for ...

Reliability Maintenance Engineer

Tustin, CA · On-site

$107K - $135K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Develop, implement, and maintain site reliability and asset management programs. * Establish ... Provide reliability engineering support for capital projects and equipment installations.

Reliability Maintenance Engineer

Tustin, CA · On-site

$107K - $135K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Develop, implement, and maintain site reliability and asset management programs. * Establish ... Provide reliability engineering support for capital projects and equipment installations.

... SRE observability practices. • Continuously expand monitoring coverage to reduce blind spots and improve operational intelligence. • Automate and maintain identity lifecycle workflows using ...

Software Engineering Manager

Irvine, CA · On-site +1

  • Medical

  • Dental

  • Vision

  • Life

  • PTO

... on SRE instincts for deployment, monitoring, telemetry, observability, and recovery. The role will partner with product engineering teams, infrastructure and SRE, and developer experience and AI ...

Software Engineer, DevOps

Irvine, CA · On-site

$56 - $76.75/hr

... observability for distributed robotics and AI systems. • Collaborate cross-functionally with ... SRE, or related roles. • Strong experience with Kubernetes, Docker, and cloud-native ...

Showing results 41-60

Observability Site Reliability Engineer information

See Riverside, CA salary details

$11

$66

$95

How much do observability site reliability engineer jobs pay per hour?

As of Aug 16, 2026, the average hourly pay for observability site reliability engineer in Riverside, CA is $66.50, according to ZipRecruiter salary data. Most workers in this role earn between $57.16 and $76.01 per hour, depending on experience, location, and employer.

What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?

AspectObservability Site Reliability EngineerMonitoring Engineer
FocusEnsuring system reliability through observability, automation, and incident responseImplementing and managing monitoring tools and dashboards
SkillsCloud platforms, scripting, incident management, observability toolsMonitoring tools, alerting systems, data analysis
Work EnvironmentDevOps teams, cloud infrastructure, large-scale systemsOperations teams, infrastructure monitoring

While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.

What are popular job titles related to Observability Site Reliability Engineer jobs in Riverside, CA?

For Observability Site Reliability Engineer jobs in Riverside, CA, the most frequently searched job titles are:

What job categories do people searching Observability Site Reliability Engineer jobs in Riverside, CA look for?

The top searched job categories for Observability Site Reliability Engineer jobs in Riverside, CA are:

What cities near Riverside, CA are hiring for Observability Site Reliability Engineer jobs?

Cities near Riverside, CA with the most Observability Site Reliability Engineer job openings:

Staff AI Platform & Reliability Engineer

eJam, Inc.

Santa Ana, CA • On-site

$172 - $220/hr

Other

Medical, Dental, Vision, Retirement, PTO

Posted 16 days ago


Job description

Location: Santa Ana, California — on‑site
Employment type: Full‑time
Salary: $172,000 – $220,000

About us

Created.ai is an AI‑powered creative platform combining text, image and video generation into production tools for commercial creative work. The platform is operated by eJam, a multi‑brand consumer products and technology company.

We are expanding our engineering team with two appointments covering our AI platform infrastructure and our product engineering surface. Both roles carry substantial technical ownership and report directly into engineering leadership.

This position is AI‑native by design. Our engineers work with agentic coding tools as a primary part of their daily workflow, which allows a small team to operate at a scope that would conventionally require a much larger one. We are looking for engineers whose judgement, architectural thinking and review discipline scale that leverage rather than being replaced by it.

Role summary

We are seeking a Staff Engineer to take ownership of the AI generation services, provider integrations, billing integrity, tenant security and overall platform reliability underpinning created.ai. The role is Python and GCP‑first, with sufficient TypeScript proficiency to trace and modify cross‑service contracts.

This is a senior individual contributor position with architectural authority over the platform layer.

Key responsibilities

  • Own the design, delivery and operation of AI generation services and all third‑party provider integrations
  • Ensure billing and usage‑metering integrity across the platform, including reconciliation and credit accounting
  • Maintain tenant isolation and platform security controls
  • Lead reliability engineering: failure handling, degradation strategy, capacity and incident response
  • Strengthen deployment safeguards, observability and cost attribution
  • Set technical standards for the Python services and mentor engineers working within them
  • Python 3.12 with FastAPI, Pydantic, asyncio and strict typing in production
  • Google Cloud Platform: Cloud Run, Pub/Sub, Cloud Tasks, GCS, Firestore, Cloud SQL/Postgres
  • Demonstrated experience with webhooks, queues, retries, idempotency, dead‑letter queues and durable asynchronous jobs
  • SQLAlchemy and Alembic, together with billing or usage‑metering experience
  • Infrastructure and delivery tooling: Terraform, IAM/OIDC, Secret Manager, CI/CD
  • Production integrations with image, video or LLM providers
  • Working proficiency in NestJS/TypeScript sufficient to modify cross‑service contracts
  • Strong background in observability, cost tracking and incident debugging
  • Fluency with agentic coding tools (Claude Code, Codex, Cursor or comparable), including the ability to scope work for them, review their output critically and maintain architectural coherence across AI‑assisted changes
  • Health, Dental, Vision
  • 401k Plan
  • PTO Plan
  • 14 observed local holidays
  • Stock options
  • Amazing, pet‑friendly office environment
  • Equipment budget and learning allowance
  • Real technical ownership within a small engineering team, and visible impact on a commercially active AI product with real users and real scale constraints
  • Direct involvement in product decisions — the team is small enough that there is nowhere to hide, in both directions
#J-18808-Ljbffr

Ejam logo

About Ejam

Sourced by ZipRecruiter

Industry

Marketing

Company size

11 - 50 Employees

Headquarters location

Newport Beach, CA, US

Year founded

2017