1

Observability Site Reliability Engineer Jobs in Riverside, CA

At least 2+ years of experience in the SRE space, with a focus on observability and automation. * Familiarity with SRE core principles (e.g. SLO, SLA, SLI, Error Budget, etc.). * Hands‑on ...

Site Reliability Engineer III

Irvine, CA · On-site

$61.25 - $81.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank ... Experience in observability such as white and black box monitoring, service level objective ...

Site Reliability Engineer III

Irvine, CA · On-site

$61.25 - $81.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank ... Experience in observability such as white and black box monitoring, service level objective ...

Site Reliability Engineer III

Irvine, CA · On-site

$61.25 - $81.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Commercial and Investment Bank ... Experience in observability such as white and black box monitoring, service level objective ...

Site Reliability Engineer III

Irvine, CA · On-site

$117 - $146.50/hr

Fundamentally, we're looking for someone who lives and breathes observability and automation, and ... At least 2+ years of experience in the SRE space, with a focus on observability and automation

Senior Site Reliability Engineer

Irvine, CA · On-site

$60.50 - $80.25/hr

... Observability for Microservices and cloud platforms like AWS, OCI, Azure, and GCP. • Write and ... a Site Reliability Engineer. • Proficiency in programming and scripting languages like Java ...

next page

Showing results 1-20

Observability Site Reliability Engineer information

See Riverside, CA salary details

$11

$66

$95

How much do observability site reliability engineer jobs pay per hour?

As of Aug 15, 2026, the average hourly pay for observability site reliability engineer in Riverside, CA is $66.50, according to ZipRecruiter salary data. Most workers in this role earn between $57.16 and $76.01 per hour, depending on experience, location, and employer.

What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?

AspectObservability Site Reliability EngineerMonitoring Engineer
FocusEnsuring system reliability through observability, automation, and incident responseImplementing and managing monitoring tools and dashboards
SkillsCloud platforms, scripting, incident management, observability toolsMonitoring tools, alerting systems, data analysis
Work EnvironmentDevOps teams, cloud infrastructure, large-scale systemsOperations teams, infrastructure monitoring

While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.

What are popular job titles related to Observability Site Reliability Engineer jobs in Riverside, CA?

For Observability Site Reliability Engineer jobs in Riverside, CA, the most frequently searched job titles are:

What job categories do people searching Observability Site Reliability Engineer jobs in Riverside, CA look for?

The top searched job categories for Observability Site Reliability Engineer jobs in Riverside, CA are:

What cities near Riverside, CA are hiring for Observability Site Reliability Engineer jobs?

Cities near Riverside, CA with the most Observability Site Reliability Engineer job openings:

AMS / Site Reliability Engineer (SRE) Telematics & Connected Car

Info Way Solutions

Irvine, CA • On-site

$61.25 - $81.25/hr

Other

Posted 9 days ago


Job description

Job Summary

We are seeking an experienced Application Management Services (AMS) / Site Reliability Engineer (SRE) to support mission-critical applications within the Telematics and Connected Car ecosystem. The ideal candidate will have strong experience in Incident Management, Application Support, SRE practices, monitoring and observability tools, and production support for automotive connected vehicle platforms.

The candidate will be responsible for ensuring application availability, monitoring production environments, driving incident resolution, performing root cause analysis (RCA), and collaborating with cross-functional engineering teams to improve system reliability and operational excellence.


Key Responsibilities
  • Provide production support for enterprise applications within the Telematics and Connected Car ecosystem.
  • Monitor application health, system performance, and infrastructure using observability tools.
  • Manage the complete Incident Management lifecycle, including Detection, Triage, Resolution, Recovery, and Root Cause Analysis (RCA).
  • Act as the primary point of contact during production incidents and coordinate with cross-functional support teams.
  • Perform application troubleshooting, log analysis, and production issue resolution.
  • Analyze recurring issues and recommend preventive measures to improve system stability.
  • Create and maintain operational runbooks, SOPs, and knowledge base articles.
  • Collaborate with development, QA, infrastructure, cloud, and networking teams for incident resolution and deployment support.
  • Support production releases, change management activities, and post-deployment validation.
  • Drive continuous service improvements by identifying automation opportunities and improving monitoring capabilities.
  • Prepare incident reports, service health dashboards, and operational metrics.

Required Skills & Qualifications
  • Bachelor''s degree in Computer Science, Information Technology, Engineering, or a related field.
  • 8+ years of experience in Application Support, Production Support, AMS, or Site Reliability Engineering (SRE).
  • Hands-on experience working within an SRE or Incident Management team.
  • Strong understanding of the complete Incident Management Lifecycle:
    • Detection
    • Triage
    • Resolution
    • Recovery
    • Root Cause Analysis (RCA)
  • Experience with monitoring and observability tools such as:
    • Dynatrace
    • Grafana
    • ELK Stack (Elasticsearch, Logstash, Kibana)
    • Splunk
    • Prometheus
  • Experience supporting enterprise applications in production environments.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent communication and stakeholder management skills.
  • Experience working in Agile/Scrum environments.

Domain Expertise (Mandatory)
  • Experience in the Telematics or Connected Car ecosystem.
  • Understanding of Connected Vehicle architecture and production support.
  • Experience supporting vehicle connectivity platforms, telematics services, remote diagnostics, or connected mobility applications.
  • Knowledge of automotive communication concepts and connected vehicle technologies is highly preferred.

Preferred Skills
  • AWS, Azure, or Google Cloud Platform Cloud Platforms.
  • Kubernetes and Docker.
  • Linux/Unix Administration.
  • Python, Bash, or Shell Scripting.
  • CI/CD tools (Jenkins, GitHub Actions, Azure DevOps).
  • ITIL Foundation Certification.
  • ServiceNow Incident Management.
  • Automotive communication protocols (CAN, CAN FD, UDS, MQTT).
  • Knowledge of SLOs, SLIs, Error Budgets, and Reliability Engineering best practices.

Technology Environment
  • Application Management Services (AMS)
  • Site Reliability Engineering (SRE)
  • Incident Management
  • Production Support
  • Dynatrace
  • Grafana
  • ELK Stack
  • Splunk
  • Prometheus
  • ServiceNow
  • Linux / Unix
  • Python / Shell Scripting
  • AWS / Azure / Google Cloud Platform
  • Docker
  • Kubernetes
  • CI/CD
  • Git
  • Agile / Scrum
  • Telematics
  • Connected Car
  • Automotive Production Support