2

Remote Chaos Engineering Jobs in California (NOW HIRING)

Site Reliability Engineer

San Francisco, CA · Remote

$67.25 - $89.25/hr

Remote (US) Department: Cloud Platform Engineering / SRE/Reliability Position summary The Site ... Drive chaos engineering, game days, and reliability testing programs * Produce SLA performance ...

Test Automation Engineer

Sunnyvale, CA · On-site +1

$54 - $71.25/hr

Ensure resilience and stability through chaos engineering techniques. * Manage and maintain test environments, including remote hardware and cloud-based platforms. * Collaborate with development ...

[Remote] Director of DevOps

Los Angeles, CA · On-site +1

$56.75 - $77.75/hr

... remote capacity. Awarded as a "best place to work" company, our culture fosters team integrity ... chaos engineering principles and "game days" to proactively test the resilience of Convoso ...

Collaborate with product, engineering, and ingestion teams to resolve escalations and improve ... Remote Work All roles are remote unless otherwise specified in the . Review the to confirm if the ...

Remote | Full-time | Reports to: VP of Product & Engineering Apiphani is a technology-enabled ... You've taken a team from "startup chaos" to predictable delivery without importing big-company ...

Senior Data Analyst

San Francisco, CA · Remote

$101K - $127K/yr

United States (Remote) -- must be U.S.-based Type: Full-time -- Early team, high ownership Data ... Design tracking architecture and guide engineering on instrumentation (Segment). * Ensure a single ...

Remote Chaos Engineering information

What does a remote chaos engineering do?

A remote chaos engineering professional designs and executes experiments to intentionally disrupt systems in order to identify vulnerabilities and improve resilience. They use tools like Chaos Monkey or Gremlin and often work with cloud environments, monitoring system behavior to ensure reliability and fault tolerance. Strong scripting skills and understanding of distributed systems are essential for this role.

What is remote chaos engineering?

Remote Chaos Engineering is the practice of testing distributed systems' resilience by intentionally introducing failures and disruptions in remote or cloud environments. The goal is to identify weaknesses and improve system reliability by simulating real-world incidents, such as network outages or server crashes, in a controlled manner. This approach helps teams understand how their applications behave under stress and develop strategies to mitigate future incidents. Remote Chaos Engineering is particularly valuable for organizations leveraging cloud infrastructure and remote services, ensuring robust performance even under unexpected conditions.

What are some common challenges faced by professionals working in remote chaos engineering roles?

Professionals in remote chaos engineering often encounter challenges such as coordinating experiments across distributed teams, ensuring clear communication about system vulnerabilities, and managing the complexity of large-scale systems without direct, on-site access. Establishing robust monitoring and rollback procedures is essential to minimize risk during remote testing. Additionally, building trust with development and operations teams is key, as chaos engineering often involves intentionally introducing failures to improve system resilience.

Which remote chaos engineering jobs can be done remotely?

Remote chaos engineering jobs are commonly available in roles such as Site Reliability Engineer, DevOps Engineer, or SRE, which often involve designing and testing system resilience using tools like Chaos Monkey or Gremlin. These positions typically require strong scripting skills and familiarity with cloud platforms, and they can often be performed entirely remotely depending on the company's policies.

What are the key skills and qualifications needed to thrive as a remote chaos engineer?

To thrive as a Remote Chaos Engineer, you need a strong background in software engineering, systems architecture, and site reliability, often supported by a degree in computer science or a related field. Familiarity with chaos engineering platforms (such as Gremlin or Chaos Monkey), cloud environments (AWS, Azure, GCP), and automation tools is typically required. Strong problem-solving abilities, clear communication, and a collaborative mindset help you effectively identify weaknesses and drive reliability improvements across distributed teams. These skills are crucial for proactively uncovering system vulnerabilities, ensuring system resilience, and maintaining high availability in complex, remote-first infrastructures.

What is the difference between Remote Chaos Engineering vs Remote Site Reliability Engineer?

AspectRemote Chaos EngineeringRemote Site Reliability Engineer
Primary FocusDesigning and executing chaos experiments to improve system resilienceEnsuring system reliability, availability, and performance through monitoring and automation
Skills & CertificationsKnowledge of chaos engineering tools, scripting, cloud platformsMonitoring tools, scripting, cloud infrastructure, SRE certifications
Work EnvironmentCollaborates with development and operations teams, often in DevOps cultureWorks closely with engineering teams to maintain system health and SLAs

While both roles focus on system stability, Remote Chaos Engineering specializes in testing system resilience through chaos experiments, whereas Remote Site Reliability Engineers focus on maintaining overall system reliability and performance. Both roles require scripting skills and cloud knowledge, but their core objectives differ: one proactively tests, the other maintains system health.

What are the most commonly searched types of Chaos Engineering jobs in California? The most popular types of Chaos Engineering jobs in California are:
What are popular job titles related to Remote Chaos Engineering jobs in California? For Remote Chaos Engineering jobs in California, the most frequently searched job titles are:
What job categories do people searching Remote Chaos Engineering jobs in California look for? The top searched job categories for Remote Chaos Engineering jobs in California are:
What cities in California are hiring for Remote Chaos Engineering jobs? Cities in California with the most Remote Chaos Engineering job openings:

Site Reliability Engineer

STN Inc

San Francisco, CA • Remote

$67.25 - $89.25/hr

Full-time

Re-posted 24 days ago


Job description

Site Reliability Engineer

Platform and software · shared across customers

Reports to: Director, Site Reliability

Location: Remote (US)

Department: Cloud Platform Engineering / SRE/Reliability

Position summary

The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.

Key responsibilities
  • Define and operate Service Level Objectives (SLOs) aligned with customer SLAs

  • Build and maintain the observability stack including metrics, logs, traces, and alerting

  • Lead incident response and chair post-incident reviews

  • Drive automation to reduce toil and improve mean-time-to-recover (MTTR)

  • Author and maintain operational runbooks alongside the NOC

  • Manage on-call rotation, escalation paths, and incident-management tooling

  • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering

  • Drive chaos engineering, game days, and reliability testing programs

  • Produce SLA performance reports in coordination with the SLA Manager

  • Mentor junior engineers and contribute to engineering culture

Required qualifications
  • 5+ years in SRE, DevOps, or production engineering roles

  • Strong programming skills in Go, Python, or both

  • Hands-on experience operating Kubernetes-based platforms at scale

  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)

  • Strong incident management experience including major-incident command

Preferred qualifications
  • GPU or HPC platform operational experience

  • Familiarity with SLA-driven customer environments and credit calculations

  • Experience with chaos engineering tools (Gremlin, Litmus, or similar)

  • Published SRE content or contributions