1

Site Reliability Engineer Sre Jobs in Riverside, CA

Site Reliable Engineer (US - Remote)

Irvine, CA · On-site +1

$60.75 - $80.75/hr

The Site Reliability Engineer will improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions. You will combine software engineering with hands-on NOC ...

The Reliability Engineer will define, design, develop, and monitor the site's critical asset register to provide standardized processes and procedures for the proper operation and care of key ...

Reliability Engineer

Irvine, CA · On-site

$110K - $138K/yr

The Reliability Engineer will define, design, develop, and monitor the site's critical asset register to provide standardized processes and procedures for the proper operation and care of key ...

Reliability Engineer

Irvine, CA · On-site

$110 - $140/hr

The Reliability Engineer will define, design, develop, and monitor the site's critical asset register to provide standardized processes and procedures for the proper operation and care of key ...

Reliability Engineer

Irvine, CA · On-site

$110K - $138K/yr

The Reliability Engineer will define, design, develop, and monitor the site's critical asset register to provide standardized processes and procedures for the proper operation and care of key ...

Lead Data Engineer

Tustin, CA · On-site

$180K - $190K/yr

Platform Reliability & Operations * Establish and implement Site Reliability Engineering (SRE) practices. * Define and monitor platform SLAs, SLOs, and key operational metrics. * Lead platform ...

Director, Platform Engineering

Irvine, CA · On-site

$171.30 - $299.80/hr

Ensure SRE, DevOps, release, and performance functions deliver consistently against SLOs, release cadences, and quality benchmarks.* **TRANSFORM: AI & Automation Agenda -** Lead the transformation of ...

Ensure SRE, DevOps, release, and performance functions deliver consistently against SLOs, release cadences, and quality benchmarks. * TRANSFORM: AI & Automation Agenda - Lead the transformation of ...

Ensure SRE, DevOps, release, and performance functions deliver consistently against SLOs, release cadences, and quality benchmarks. * TRANSFORM: AI & Automation Agenda - Lead the transformation of ...

Showing results 21-40

Site Reliability Engineer Sre information

See Riverside, CA salary details

$11

$66

$95

How much do site reliability engineer sre jobs pay per hour?

As of Sep 6, 2026, the average hourly pay for site reliability engineer sre in Riverside, CA is $66.50, according to ZipRecruiter salary data. Most workers in this role earn between $57.16 and $76.01 per hour, depending on experience, location, and employer.

What is a site reliability engineer (SRE)?

A Site Reliability Engineer (SRE) is a professional who combines software engineering and IT operations skills to ensure reliable, scalable, and efficient systems. SREs automate system operations, monitor system health, and manage incident response to minimize downtime and ensure high availability. Their work bridges the gap between development and operations teams, focusing on building robust infrastructure and processes. SREs often use tools and practices like automation, monitoring, and performance tuning to maintain service reliability. They also set and measure service level objectives (SLOs) to align system performance with business goals.

What are the key skills and qualifications needed to thrive as a site reliability engineer (SRE), and why are they important?

To thrive as a Site Reliability Engineer (SRE), you need a solid background in software engineering, systems administration, and automation, often supported by a degree in computer science or a related field. Familiarity with cloud platforms (like AWS, GCP, or Azure), containerization (Docker, Kubernetes), and monitoring tools (Prometheus, Grafana), as well as scripting languages such as Python or Bash, is typically required. Strong problem-solving, communication, and collaboration skills help SREs manage incidents and work effectively with development and operations teams. These competencies are essential to ensure system reliability, optimize performance, and maintain seamless service availability.

What are some of the most common challenges faced by site reliability engineers (SREs), and how can they be addressed?

Site Reliability Engineers often face challenges such as managing the balance between reliability and rapid feature deployment, handling on-call responsibilities, and automating manual processes. To address these, SREs work closely with development and operations teams to implement strong monitoring, establish clear service level objectives (SLOs), and continually improve incident response procedures. Building a culture of blameless postmortems and investing in automation can also help reduce repetitive work and improve overall system reliability.

What is the difference between Site Reliability Engineer Sre vs DevOps Engineer?

AspectSite Reliability Engineer SreDevOps Engineer
CredentialsTypically requires experience in systems engineering, scripting, and monitoring toolsOften has certifications in cloud platforms, automation, and CI/CD tools
Work EnvironmentFocuses on maintaining system reliability, scalability, and incident responseEmphasizes automation, deployment, and continuous integration/delivery
Industry UsageCommon in tech, finance, and large-scale cloud servicesWidely used across startups, tech companies, and enterprises

While both roles aim to improve system performance and automation, Site Reliability Engineers Sre primarily focus on reliability and incident management, whereas DevOps Engineers concentrate on deployment automation and continuous integration. The roles often overlap but serve distinct core functions within IT and software development teams.

What are popular job titles related to Site Reliability Engineer Sre jobs in Riverside, CA?

For Site Reliability Engineer Sre jobs in Riverside, CA, the most frequently searched job titles are:

What job categories do people searching Site Reliability Engineer Sre jobs in Riverside, CA look for?

The top searched job categories for Site Reliability Engineer Sre jobs in Riverside, CA are:

What cities near Riverside, CA are hiring for Site Reliability Engineer Sre jobs?

Cities near Riverside, CA with the most Site Reliability Engineer Sre job openings:

Infographic showing various Site Reliability Engineer Sre job openings in Riverside, CA as of August 2026, with employment types broken down into 1% As Needed, 78% Full Time, 17% Part Time, 1% Temporary, 2% Contract, and 1% Nights. Highlights an 93% Physical, 2% Hybrid, and 5% Remote job distribution, with an average salary of $138,320 per year, or $66.5 per hour.

Site Reliable Engineer (US - Remote)

AXON-Networks

Irvine, CA • On-site, Remote

$60.75 - $80.75/hr

Full-time

Posted 7 days ago


Job description

AXON Networks delivers a robust AI-driven, analytics-based orchestration platform and a wide portfolio of next-gen high-speed routers that leverage the newest Wi-Fi technologies. Together, these technologies give ISPs the ability to manage and troubleshoot their networks in real time, and to deliver an outstanding customer experience.
AXON Networks is a trusted strategic partner for its customers, helping them evaluate their current technologies and business models, and creating and executing strategies that enable them to innovate faster, accelerate their digital transformations, and strengthen their relationships with consumers.
AXON Networks is headquartered in Irvine, CA USA with Asia HQ in Singapore and also operating in Denmark, Spain and Vietnam.
The Site Reliability Engineer will improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions. You will combine software engineering with hands-on NOC operations to make the complete cloud-to-device service path observable, supportable and resilient at fleet scale.
You will help establish practical SRE capabilities inside the NOC while partnering closely with Support, Operations, cloud and DevOps Engineering. You will participate in a sustainable on-call rotation and improve the NOC's ability to diagnose customer-impacting issues.
Role mandate
  • Own reliability outcomes for assigned cloud services

  • Improve observability, capacity, resilience and recovery

  • Define and operationalize service-level indicators, service-level objectives and actionable alerting.

  • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes.

  • Lead technically during incidents, drive evidence-based learning and ensure high-value corrective actions are completed.

What you will own
  • Establish reliability baselines, SLIs, SLOs and error budgets for cloud services and critical device-management workflows such as onboarding, provisioning, configuration, telemetry collection, command execution and firmware delivery.

  • Trace failures across the end-to-end service path: cloud APIs and microservices, Kubernetes and infrastructure, databases and messaging, internet and access-network dependencies, device-management protocols and the devices

  • Identify fleet-wide and customer-specific failure patterns involving device reachability, session stability, configuration drift, command latency, telemetry gaps, firmware behavior and cloud capacity.

  • Contribute operability requirements and production evidence during design and readiness reviews

  • Maintain NOC dashboards for service health, device reachability, provisioning success, command and telemetry performance, firmware adoption and customer impact.

  • Participate in the NOC production on-call rotation and serve as a technical incident lead or senior troubleshooter when appropriate.

  • Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging platforms, device-management sessions and CPE behavior.

  • Coordinate evidence gathering and technical escalation with service-provider customers, Engineering, firmware, DevOps and vendors while maintaining clear mitigation, recovery and handoff.

  • Lead or contribute to post-incident reviews; convert recurring device, platform and process failures into prioritized and measurable corrective actions.

  • Develop production-grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery.

  • Improve CI/CD and GitOps practices for operational software and infrastructure, including automated testing, release validation, progressive delivery and rollback readiness.

  • Manage or contribute to infrastructure as code, configuration as code and reusable self-service patterns for cloud and NOC operations.

  • Measure NOC toil and partner with Automation & Tools Engineers to prioritize durable platform capabilities instead of fragmented one-off scripts.

  • Develop capacity models for service-provider growth, managed-device populations, telemetry volume, messaging throughput, API demand and rollout events.

  • Create and maintain runbooks, troubleshooting decision trees, service maps, device and cloud dependency records, known-error guidance and operational knowledge.

  • Coach NOC and Support personnel on diagnosis, safe mitigation, evidence capture and escalation across cloud, network and CPE layers.

  • Build self-service diagnostic views and tools that help the NOC determine scope, affected customers, device cohorts, likely fault domain and next action.

  • Share reliability insights with Engineering and Product and contribute to reliability reviews, operational-readiness reviews and continuous-improvement priorities.

Required qualifications
  • 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role.

  • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code.

  • Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers.

  • Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments.

  • Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/CD and policy-as-code.

  • Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis.

  • Strong troubleshooting & debugging skills in Kubernetes platforms.

  • Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring.

  • Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems.

  • Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices.

  • Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents.

  • Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning.

  • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet- or session-level troubleshooting.

  • Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams.

  • Bachelor's degree in computer science, engineering or equivalent practical experience.

Preferred qualifications
  • Experience supporting cloud-managed CPEs such as broadband gateways, routers, ONTs, Wi-Fi/mesh systems or similar edge devices in a service-provider environment.

  • Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management.

  • Experience supporting messaging and streaming platforms such as Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes.

  • Understanding of access technologies such as GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless and how CPE, ONTs and provider networks interact.

  • Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices.

  • Experience building auto-remediation, safe self-service operations or internal reliability platforms.

  • Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment.

This position is fully remote within North America. Please note that we are unable to offer visa sponsorship for this role.
Annual salary range: $160,000 - $200,000
Join AXON Networks!
At AXON Networks, we promote equal opportunities in all our recruitment processes, ensuring non-discrimination on the basis of gender, age, origin, disability, or any other personal circumstances. We assess talent based on objective criteria and foster an inclusive and diverse working environment.
axon-networks.com
axon-networks.hire.trakstar.com