2

Remote Chaos Engineering Jobs (NOW HIRING)

Site Reliability Engineer

San Francisco, CA · Remote

$67.25 - $89.25/hr

Remote (US) Department: Cloud Platform Engineering / SRE/Reliability Position summary The Site ... Drive chaos engineering, game days, and reliability testing programs * Produce SLA performance ...

Data Scientist, AI/ML

$220K - $290K/yr

As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the ... Working in Remote first environments * Has been on-call and participated in an incident management ...

Senior Backend Engineer

$220K - $290K/yr

... Chaos Engineering tooling * Leverage strong collaboration and communication skills to deliver new features within a remote culture * Partner with product and other business units to understand ...

$94K - $124K/yr

Implement chaos engineering practices by designing fault-injection experiments that proactively ... Fully remote work environment with global collaboration. * Opportunity to work with modern cloud ...

Lead the organization from reactive firefighting to a predictive, self-healing culture through the aggressive adoption of Chaos Engineering and AlOps Job Designation Remote: Employee is not required ...

Lead the organization from reactive firefighting to a predictive, self-healing culture through the aggressive adoption of Chaos Engineering and AlOps Job Designation Remote: Employee is not required ...

DevOps & SRE Engineer

Apex, NC · Remote

$100K - $150K/yr

Job Title: DevOps & SRE Engineer Location: 100% Remote (U.S.) Position Type: Full-time, Direct W2 ... Exposure to chaos engineering practices and tools such as Chaos Monkey, Gremlin, or Litmus. * Hands ...

Shift the team from reactive incident response to predictive mitigation by implementing Chaos Engineering practices and automated self-healing mechanisms Job Designation Remote: Employee is not ...

Infrastructure / DevOps Engineer

Birmingham, AL · On-site +1

$49.50 - $67.75/hr

... chaos engineering, and reliability reviews Location Remote (US). If you're in the Birmingham, AL area, this role is hybrid from our office (3 days in office / 2 remote). Requirements * 5+ years in ...

Infrastructure / DevOps Engineer

Birmingham, AL · Remote

$54 - $74/hr

... chaos engineering, and reliability reviews Location Remote (US). If you're in the Birmingham, AL area, this role is hybrid from our office (3 days in office / 2 remote). Requirements * 5+ years in ...

Remote WHO WE ARE AND WHAT WE DO Mosaic Clinical Technologies™ is pioneering a new imaging ... Drive disaster recovery, chaos engineering, load testing, and resiliency initiatives to improve ...

Infrastructure / DevOps Engineer

Birmingham, AL · On-site +1

$49.50 - $67.75/hr

... chaos engineering, and reliability reviews Location Remote (US). If you're in the Birmingham, AL area, this role is hybrid from our office (3 days in office / 2 remote). Requirements * 5+ years in ...

next page

Showing results 1-20

Remote Chaos Engineering information

See salary details

$73K

$194.7K

$254K

How much do remote chaos engineering jobs pay per year?

As of Aug 11, 2026, the average yearly pay for remote chaos engineering in the United States is $194,709.00, according to ZipRecruiter salary data. Most workers in this role earn between $141,500.00 and $253,000.00 per year, depending on experience, location, and employer.

What does a remote chaos engineering do?

A remote chaos engineering professional designs and executes experiments to intentionally disrupt systems in order to identify vulnerabilities and improve resilience. They use tools like Chaos Monkey or Gremlin and often work with cloud environments, monitoring system behavior to ensure reliability and fault tolerance. Strong scripting skills and understanding of distributed systems are essential for this role.

What is remote chaos engineering?

Remote Chaos Engineering is the practice of testing distributed systems' resilience by intentionally introducing failures and disruptions in remote or cloud environments. The goal is to identify weaknesses and improve system reliability by simulating real-world incidents, such as network outages or server crashes, in a controlled manner. This approach helps teams understand how their applications behave under stress and develop strategies to mitigate future incidents. Remote Chaos Engineering is particularly valuable for organizations leveraging cloud infrastructure and remote services, ensuring robust performance even under unexpected conditions.

What are some common challenges faced by professionals working in remote chaos engineering roles?

Professionals in remote chaos engineering often encounter challenges such as coordinating experiments across distributed teams, ensuring clear communication about system vulnerabilities, and managing the complexity of large-scale systems without direct, on-site access. Establishing robust monitoring and rollback procedures is essential to minimize risk during remote testing. Additionally, building trust with development and operations teams is key, as chaos engineering often involves intentionally introducing failures to improve system resilience.

Which remote chaos engineering jobs can be done remotely?

Remote chaos engineering jobs are commonly available in roles such as Site Reliability Engineer, DevOps Engineer, or SRE, which often involve designing and testing system resilience using tools like Chaos Monkey or Gremlin. These positions typically require strong scripting skills and familiarity with cloud platforms, and they can often be performed entirely remotely depending on the company's policies.

What are the key skills and qualifications needed to thrive as a remote chaos engineer?

To thrive as a Remote Chaos Engineer, you need a strong background in software engineering, systems architecture, and site reliability, often supported by a degree in computer science or a related field. Familiarity with chaos engineering platforms (such as Gremlin or Chaos Monkey), cloud environments (AWS, Azure, GCP), and automation tools is typically required. Strong problem-solving abilities, clear communication, and a collaborative mindset help you effectively identify weaknesses and drive reliability improvements across distributed teams. These skills are crucial for proactively uncovering system vulnerabilities, ensuring system resilience, and maintaining high availability in complex, remote-first infrastructures.

What is the difference between Remote Chaos Engineering vs Remote Site Reliability Engineer?

AspectRemote Chaos EngineeringRemote Site Reliability Engineer
Primary FocusDesigning and executing chaos experiments to improve system resilienceEnsuring system reliability, availability, and performance through monitoring and automation
Skills & CertificationsKnowledge of chaos engineering tools, scripting, cloud platformsMonitoring tools, scripting, cloud infrastructure, SRE certifications
Work EnvironmentCollaborates with development and operations teams, often in DevOps cultureWorks closely with engineering teams to maintain system health and SLAs

While both roles focus on system stability, Remote Chaos Engineering specializes in testing system resilience through chaos experiments, whereas Remote Site Reliability Engineers focus on maintaining overall system reliability and performance. Both roles require scripting skills and cloud knowledge, but their core objectives differ: one proactively tests, the other maintains system health.

More about Remote Chaos Engineering jobs
What cities are hiring for Remote Chaos Engineering jobs? Cities with the most Remote Chaos Engineering job openings:
What are the most commonly searched types of Chaos Engineering jobs? The most popular types of Chaos Engineering jobs are:
What states have the most Remote Chaos Engineering jobs? States with the most job openings for Remote Chaos Engineering jobs include:
Infographic showing various Remote Chaos Engineering job openings in the United States as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% Remote job distribution, with an average salary of $194,709 per year, or $93.6 per hour.

Manager, Software Engineering (Resilience Engineering)

Affirm

Remote

Full-time

Medical, Dental, Vision

Re-posted 12 days ago


Job description

Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest.
We are seeking a seasoned Engineering Manager to lead our Resilience Engineering team. This role is critical in ensuring the safety and reliability of our production systems through proactive validation techniques, including production load testing and chaos engineering.
You will lead the development of systems and practices that allow engineers to safely test system behavior under stress and failure conditions in production, ensuring issues are discovered and mitigated before they impact real users.
What you'll do
Leadership & Strategy
  • Define and drive the vision for resilience engineering at Affirm, with a focus on production load testing and chaos engineering as first-class engineering practices.
  • Lead and mentor a team of engineers building platforms and tooling for safe production experimentation.
  • Partner with infrastructure, product, and security leadership to embed resilience validation into the software development lifecycle.
  • Establish best practices for safely testing system limits and failure scenarios in production.

Systems & Operations
  • Own the design and evolution of platforms that enable safe, controlled production load testing and fault injection.
  • Ensure strong safeguards are in place, including isolation boundaries, approval workflows, and automated rollback mechanisms to protect real users.
  • Build systems that provide end-to-end observability, traceability, and auditability for all resilience experiments.
  • Drive reliability improvements by systematically identifying weaknesses through load testing and chaos experiments.
  • Establish monitoring, alerting, and incident response practices tailored to proactive resilience validation.

Collaboration & Enablement
  • Work closely with engineering teams to design and execute production load tests and chaos experiments safely.
  • Partner with infrastructure teams to build guardrails around tests and experimentations.
  • Enable teams to adopt resilience practices by providing reusable tooling, frameworks, and standardized workflows.
  • Identify systemic weaknesses and lead cross-functional efforts to improve reliability and fault tolerance.
  • Evangelize a culture of "test failure before failure tests you" across the organization.

What we look for
  • Proven experience leading engineering teams in reliability, infrastructure, or distributed systems.
  • Hands-on experience with production load testing, chaos engineering, or large-scale system validation.
  • Experience with leveraging a chaos engineering vendor such as Gremlin, Harness, or something similar.
  • Strong understanding of failure modes in distributed systems, including latency, partial failure, and cascading outages.
  • Experience building or operating systems with strong safety guarantees (isolation, rate limiting, guardrails, auditability).
  • Familiarity with cloud-native environments (AWS, Kubernetes) and observability tooling.
  • Strong programming background (e.g., Python, Kotlin, Java, or similar).
  • Excellent problem-solving skills and the ability to balance long-term resilience investments with immediate business needs.
  • Strong communication and leadership skills, with a track record of influencing engineering practices across teams.
  • This position requires either equivalent practical experience or a Bachelor's degree in a related field.

Base Pay Grade - P
Equity Grade - 13
Employees new to Affirm typically come in at the start of the pay range. Affirm focuses on providing a simple and transparent pay structure which is based on a variety of factors, including location, experience and job-related skills.
Base pay is part of a total compensation package that may include equity rewards, monthly stipends for health, wellness and tech spending, and benefits (including 100% subsidized medical coverage, dental and vision for you and your dependents.)
USA base pay range (CA, WA, NY, NJ, CT) per year: 230,000 - 290,000
USA base pay range (all other U.S. states) per year: 204,000 - 264,000
#LI-Remote
Affirm is proud to be a remote-first company! The majority of our roles are remote and you can work almost anywhere within the country of employment. Affirmers in proximal roles have the flexibility to work remotely, but will occasionally be required to work out of their assigned Affirm office. A limited number of roles remain office-based due to the nature of their job responsibilities.
We're extremely proud to offer competitive benefits that are anchored to our core value of people come first. Some key highlights of our benefits package include:
  • Health care coverage - Affirm covers all premiums for all levels of coverage for you and your dependents
  • Flexible Spending Wallets - generous stipends for spending on Technology, Food, various Lifestyle needs, and family forming expenses
  • Time off - competitive vacation and holiday schedules allowing you to take time off to rest and recharge
  • ESPP - An employee stock purchase plan enabling you to buy shares of Affirm at a discount

We believe It's On Us to provide an inclusive interview experience for all, including people with disabilities. We are happy to provide reasonable accommodations to candidates in need of individualized support during the hiring process.
[For U.S. positions that could be performed in Los Angeles or San Francisco] Pursuant to the San Francisco Fair Chance Ordinance and Los Angeles Fair Chance Initiative for Hiring Ordinance, Affirm will consider for employment qualified applicants with arrest and conviction records.
By clicking "Submit Application," you acknowledge that you have read Affirm's Global Candidate Privacy Notice and hereby freely and unambiguously give informed consent to the collection, processing, use, and storage of your personal information as described therein.