2

Remote Site Reliability Engineer Intern Jobs in California

Site Reliability Engineer

San Francisco, CA · Remote

$67.25 - $89.25/hr

Remote (US) Department: Cloud Platform Engineering / SRE/Reliability Position summary The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One ...

Senior Site Reliability Engineer

San Diego, CA · Remote

$60.50 - $80.50/hr

This is a remote, contract opportunity for a project Arctiq is delivering for a client. Candidates ... The Senior Site Reliability Engineer is a technical leader responsible for architecting the ...

Site Reliability Engineer

Palo Alto, CA · On-site +1

$165K - $190K/yr

About the DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-performing production systems. We work closely with ...

About the job We\'re looking for a Site Reliability Engineer (E3) to help build and operate the ... This position is remote and requires working East Coast business hours (EST). What you\'ll do

$98K - $138K/yr

The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining ... Wellness initiatives #BI-Remote Internal Employees - R365 is committed to growing talent from ...

$98K - $138K/yr

The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining ... Wellness initiatives #BI-Remote Internal Employees - R365 is committed to growing talent from ...

Senior Site Reliability Engineer

Glendale, CA · On-site +1

$60.50 - $80.25/hr

Our Site Reliability and Infrastructure Engineering team centralizes the concerns of measurement and guidance so every engineer can improve availability and efficiency in their own area of the ...

next page

Showing results 1-20

Remote Site Reliability Engineer Intern information

What is the difference between Remote Site Reliability Engineer Intern vs Remote DevOps Engineer Intern?

AspectRemote Site Reliability Engineer InternRemote DevOps Engineer Intern
FocusEnsuring system reliability, uptime, and performanceAutomating deployment, integration, and infrastructure management
SkillsMonitoring, incident response, scripting, cloud platformsCI/CD, automation tools, scripting, cloud services
Work EnvironmentCollaborates with SRE and operations teamsWorks with development and operations teams
CertificationsOften preferred: Linux, cloud certificationsOften preferred: Linux, cloud, and automation certifications

While both roles involve cloud and scripting skills, the Remote Site Reliability Engineer Intern focuses on system stability and uptime, whereas the Remote DevOps Engineer Intern emphasizes automation and deployment processes. Understanding these differences helps candidates target their skills and career goals effectively.

What job categories do people searching Remote Site Reliability Engineer Intern jobs in California look for?

The top searched job categories for Remote Site Reliability Engineer Intern jobs in California are:

What cities in California are hiring for Remote Site Reliability Engineer Intern jobs?

Cities in California with the most Remote Site Reliability Engineer Intern job openings:

Infographic showing various Remote Site Reliability Engineer Intern job openings in California as of August 2026, with employment types broken down into 68% Full Time, 19% Part Time, and 13% Contract. Highlights an 100% Remote job distribution.

Site Reliability Engineer

STN Inc

San Francisco, CA • Remote

$67.25 - $89.25/hr

Full-time

This job post has expired 1 day ago. Applications are no longer accepted.


Job description

Site Reliability Engineer

Platform and software · shared across customers

Reports to: Director, Site Reliability

Location: Remote (US)

Department: Cloud Platform Engineering / SRE/Reliability

Position summary

The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.

Key responsibilities
  • Define and operate Service Level Objectives (SLOs) aligned with customer SLAs

  • Build and maintain the observability stack including metrics, logs, traces, and alerting

  • Lead incident response and chair post-incident reviews

  • Drive automation to reduce toil and improve mean-time-to-recover (MTTR)

  • Author and maintain operational runbooks alongside the NOC

  • Manage on-call rotation, escalation paths, and incident-management tooling

  • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering

  • Drive chaos engineering, game days, and reliability testing programs

  • Produce SLA performance reports in coordination with the SLA Manager

  • Mentor junior engineers and contribute to engineering culture

Required qualifications
  • 5+ years in SRE, DevOps, or production engineering roles

  • Strong programming skills in Go, Python, or both

  • Hands-on experience operating Kubernetes-based platforms at scale

  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)

  • Strong incident management experience including major-incident command

Preferred qualifications
  • GPU or HPC platform operational experience

  • Familiarity with SLA-driven customer environments and credit calculations

  • Experience with chaos engineering tools (Gremlin, Litmus, or similar)

  • Published SRE content or contributions