2

Remote Chaos Engineering Jobs in California (NOW HIRING)

Test Automation Engineer

Sunnyvale, CA · On-site +1

$54 - $71.25/hr

  • Medical

  • Dental

  • Life

  • Retirement

  • PTO

Ensure resilience and stability through chaos engineering techniques. * Manage and maintain test environments, including remote hardware and cloud-based platforms. * Collaborate with development ...

[Remote] Director of DevOps

Los Angeles, CA · On-site +1

$56.75 - $77.75/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... remote capacity. Awarded as a "best place to work" company, our culture fosters team integrity ... chaos engineering principles and "game days" to proactively test the resilience of Convoso ...

Customer Centric Engineer

Palo Alto, CA · On-site +1

  • Medical

  • Dental

  • Vision

  • Life

  • PTO

Collaborate with product, engineering, and ingestion teams to resolve escalations and improve ... Remote Work All roles are remote unless otherwise specified in the . Review the to confirm if the ...

Remote | Full-time | Reports to: VP of Product & Engineering Apiphani is a technology-enabled ... You've taken a team from "startup chaos" to predictable delivery without importing big-company ...

Backend Software Engineer - Web/ML Dev Ops

San Mateo, CA · Remote

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... remote role. You'll build and own the backend systems that power our Computer Vision AI platform ... Work closely with Solutions Engineering to translate customer requirements into scalable backend ...

Revenue Operations Lead

San Francisco, CA · On-site +1

$179K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

USA Remote / San Francisco, CA • Full-Time About Andromeda Andromeda Cluster is a $1.5B company ... research and engineering. The Role This is a chance to join Andromeda as our first Revenue ...

Remote Chaos Engineering information

What does a remote chaos engineering do?

A remote chaos engineering professional designs and executes experiments to intentionally disrupt systems in order to identify vulnerabilities and improve resilience. They use tools like Chaos Monkey or Gremlin and often work with cloud environments, monitoring system behavior to ensure reliability and fault tolerance. Strong scripting skills and understanding of distributed systems are essential for this role.

What is remote chaos engineering?

Remote Chaos Engineering is the practice of testing distributed systems' resilience by intentionally introducing failures and disruptions in remote or cloud environments. The goal is to identify weaknesses and improve system reliability by simulating real-world incidents, such as network outages or server crashes, in a controlled manner. This approach helps teams understand how their applications behave under stress and develop strategies to mitigate future incidents. Remote Chaos Engineering is particularly valuable for organizations leveraging cloud infrastructure and remote services, ensuring robust performance even under unexpected conditions.

What are some common challenges faced by professionals working in remote chaos engineering roles?

Professionals in remote chaos engineering often encounter challenges such as coordinating experiments across distributed teams, ensuring clear communication about system vulnerabilities, and managing the complexity of large-scale systems without direct, on-site access. Establishing robust monitoring and rollback procedures is essential to minimize risk during remote testing. Additionally, building trust with development and operations teams is key, as chaos engineering often involves intentionally introducing failures to improve system resilience.

Which remote chaos engineering jobs can be done remotely?

Remote chaos engineering jobs are commonly available in roles such as Site Reliability Engineer, DevOps Engineer, or SRE, which often involve designing and testing system resilience using tools like Chaos Monkey or Gremlin. These positions typically require strong scripting skills and familiarity with cloud platforms, and they can often be performed entirely remotely depending on the company's policies.

What are the key skills and qualifications needed to thrive as a remote chaos engineer?

To thrive as a Remote Chaos Engineer, you need a strong background in software engineering, systems architecture, and site reliability, often supported by a degree in computer science or a related field. Familiarity with chaos engineering platforms (such as Gremlin or Chaos Monkey), cloud environments (AWS, Azure, GCP), and automation tools is typically required. Strong problem-solving abilities, clear communication, and a collaborative mindset help you effectively identify weaknesses and drive reliability improvements across distributed teams. These skills are crucial for proactively uncovering system vulnerabilities, ensuring system resilience, and maintaining high availability in complex, remote-first infrastructures.

What is the difference between Remote Chaos Engineering vs Remote Site Reliability Engineer?

AspectRemote Chaos EngineeringRemote Site Reliability Engineer
Primary FocusDesigning and executing chaos experiments to improve system resilienceEnsuring system reliability, availability, and performance through monitoring and automation
Skills & CertificationsKnowledge of chaos engineering tools, scripting, cloud platformsMonitoring tools, scripting, cloud infrastructure, SRE certifications
Work EnvironmentCollaborates with development and operations teams, often in DevOps cultureWorks closely with engineering teams to maintain system health and SLAs

While both roles focus on system stability, Remote Chaos Engineering specializes in testing system resilience through chaos experiments, whereas Remote Site Reliability Engineers focus on maintaining overall system reliability and performance. Both roles require scripting skills and cloud knowledge, but their core objectives differ: one proactively tests, the other maintains system health.

What are the most commonly searched types of Chaos Engineering jobs in California? The most popular types of Chaos Engineering jobs in California are:
What job categories do people searching Remote Chaos Engineering jobs in California look for? The top searched job categories for Remote Chaos Engineering jobs in California are:
What cities in California are hiring for Remote Chaos Engineering jobs? Cities in California with the most Remote Chaos Engineering job openings:

Senior Software Engineer, Observability

Together AI

San Francisco, CA • On-site, Remote

$200K - $280K/yr

Full-time

Medical

Re-posted 5 days ago


Job description

About the Role

Together AI is building the AI Acceleration Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art AI cloud infrastructure.

The AI Infrastructure team at Together AI is at the forefront of building and scaling the foundational systems that power our generative AI platform. The storage and observability team is crucial for designing, implementing, and maintaining robust distributed storage solutions, ensuring seamless data access and management. They are also responsible for developing comprehensive observability platforms, providing critical insights into system performance and GPU utilization, and proactively identifying and resolving issues.

Responsibilities

  • Design and implement a scalable observability platform (metrics, logs, traces) using tools like Prometheus, Grafana, ClickHouse, ClickStack, and OpenTelemetry, including telemetry data pipelines and log aggregation workflows.
  • Develop automated monitoring, alerting, and anomaly detection systems, including SLIs/SLOs, runbooks, and predictive analytics for critical services.
  • Build and deploy custom observability tools and infrastructure-as-code using Go, Python, Terraform, Ansible, and Helm.
  • Collaborate with engineering teams to enhance distributed tracing and application monitoring, and lead incident response with post-mortem analysis.
  • Define observability best practices.

Requirements

  • Expertise in observability platforms (Prometheus, Grafana, ClickStack, OpenTelemetry) and cloud-native monitoring services (AWS, GCP, Azure).
  • Strong programming skills in Go, Python, or similar languages, with proficiency in infrastructure-as-code tools (Terraform, Ansible, Helm).
  • Experience designing, operating, and scaling large-scale distributed systems and pipelines for high-volume data ingestion and real-time querying.
  • Deep understanding of containerization (Docker) and orchestration (Kubernetes).
  • Knowledge of microservices architecture, service mesh technologies, CI/CD pipelines, and GitOps workflows.
  • Expertise in managing databases (PostgreSQL, MongoDB, Redis) and time-series databases with high-cardinality data.

Preferred

  • Experience monitoring AI/ML infrastructure, GPU clusters, and custom metrics for model performance and training pipelines.
  • Background in high-frequency, low-latency systems monitoring, chaos engineering, and reliability testing.
  • Contributions to open-source observability projects.
  • Familiarity with security monitoring and compliance frameworks.
Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $200,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at https://www.together.ai/privacy