2

Remote Chaos Engineering Jobs (NOW HIRING)

Data Scientist, AI/ML

$220K - $290K/yr

As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the ... Working in Remote first environments * Has been on-call and participated in an incident management ...

Data Scientist, AI/ML

$220K - $290K/yr

As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the ... Working in Remote first environments * Has been on-call and participated in an incident management ...

Senior Backend Engineer

$220K - $290K/yr

... Chaos Engineering tooling * Leverage strong collaboration and communication skills to deliver new features within a remote culture * Partner with product and other business units to understand ...

Senior Backend Engineer

$220K - $290K/yr

... Chaos Engineering tooling * Leverage strong collaboration and communication skills to deliver new features within a remote culture * Partner with product and other business units to understand ...

Lead the organization from reactive firefighting to a predictive, self-healing culture through the aggressive adoption of Chaos Engineering and AlOps Job Designation Remote: Employee is not required ...

Lead the organization from reactive firefighting to a predictive, self-healing culture through the aggressive adoption of Chaos Engineering and AlOps Job Designation Remote: Employee is not required ...

... chaos engineering techniques such as edge cases, failure modes and design review • Advise on capacity planning and provide continuous assessments on systems behavior and consumption • Work with ...

... CA or Remote reporting to Production Engineering in the Cloud Infrastructure & Operations ... Experience with chaos engineering and disaster recovery planning at scale * Expertise in global ...

... CA or Remote reporting to Production Engineering in the Cloud Infrastructure & Operations ... Experience with chaos engineering and disaster recovery planning at scale * Expertise in global ...

... CA or Remote reporting to Production Engineering in the Cloud Infrastructure & Operations ... Experience with chaos engineering and disaster recovery planning at scale * Expertise in global ...

... CA or Remote reporting to Production Engineering in the Cloud Infrastructure & Operations ... Experience with chaos engineering and disaster recovery planning at scale * Expertise in global ...

Shift the team from reactive incident response to predictive mitigation by implementing Chaos Engineering practices and automated self-healing mechanisms Job Designation Remote: Employee is not ...

Infrastructure / DevOps Engineer

Birmingham, AL · Remote

$54 - $74/hr

... chaos engineering, and reliability reviews Location Remote (US). If you're in the Birmingham, AL area, this role is hybrid from our office (3 days in office / 2 remote). Requirements * 5+ years in ...

Pre/Post Sales Solutions Architect

$64.50 - $85/hr

As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the ... But as a remote company, teamwork and collaboration won't happen by accident. We approach every ...

Infrastructure / DevOps Engineer

Birmingham, AL · On-site +1

$49.50 - $67.75/hr

... chaos engineering, and reliability reviews Location Remote (US). If you're in the Birmingham, AL area, this role is hybrid from our office (3 days in office / 2 remote). Requirements * 5+ years in ...

next page

Showing results 1-20

Remote Chaos Engineering information

See salary details

$73K

$194.7K

$254K

How much do remote chaos engineering jobs pay per year?

As of Sep 1, 2026, the average yearly pay for remote chaos engineering in the United States is $194,709.00, according to ZipRecruiter salary data. Most workers in this role earn between $141,500.00 and $253,000.00 per year, depending on experience, location, and employer.

What is remote chaos engineering?

Remote Chaos Engineering is the practice of testing distributed systems' resilience by intentionally introducing failures and disruptions in remote or cloud environments. The goal is to identify weaknesses and improve system reliability by simulating real-world incidents, such as network outages or server crashes, in a controlled manner. This approach helps teams understand how their applications behave under stress and develop strategies to mitigate future incidents. Remote Chaos Engineering is particularly valuable for organizations leveraging cloud infrastructure and remote services, ensuring robust performance even under unexpected conditions.

What are the key skills and qualifications needed to thrive as a remote chaos engineer?

To thrive as a Remote Chaos Engineer, you need a strong background in software engineering, systems architecture, and site reliability, often supported by a degree in computer science or a related field. Familiarity with chaos engineering platforms (such as Gremlin or Chaos Monkey), cloud environments (AWS, Azure, GCP), and automation tools is typically required. Strong problem-solving abilities, clear communication, and a collaborative mindset help you effectively identify weaknesses and drive reliability improvements across distributed teams. These skills are crucial for proactively uncovering system vulnerabilities, ensuring system resilience, and maintaining high availability in complex, remote-first infrastructures.

What are some common challenges faced by professionals working in remote chaos engineering roles?

Professionals in remote chaos engineering often encounter challenges such as coordinating experiments across distributed teams, ensuring clear communication about system vulnerabilities, and managing the complexity of large-scale systems without direct, on-site access. Establishing robust monitoring and rollback procedures is essential to minimize risk during remote testing. Additionally, building trust with development and operations teams is key, as chaos engineering often involves intentionally introducing failures to improve system resilience.

What is the difference between Remote Chaos Engineering vs Remote Site Reliability Engineer?

AspectRemote Chaos EngineeringRemote Site Reliability Engineer
Primary FocusDesigning and executing chaos experiments to improve system resilienceEnsuring system reliability, availability, and performance through monitoring and automation
Skills & CertificationsKnowledge of chaos engineering tools, scripting, cloud platformsMonitoring tools, scripting, cloud infrastructure, SRE certifications
Work EnvironmentCollaborates with development and operations teams, often in DevOps cultureWorks closely with engineering teams to maintain system health and SLAs

While both roles focus on system stability, Remote Chaos Engineering specializes in testing system resilience through chaos experiments, whereas Remote Site Reliability Engineers focus on maintaining overall system reliability and performance. Both roles require scripting skills and cloud knowledge, but their core objectives differ: one proactively tests, the other maintains system health.

More about Remote Chaos Engineering jobs

What cities are hiring for Remote Chaos Engineering jobs?

Cities with the most Remote Chaos Engineering job openings:

What are the most commonly searched types of Chaos Engineering jobs?

The most popular types of Chaos Engineering jobs are:

What states have the most Remote Chaos Engineering jobs?

States with the most job openings for Remote Chaos Engineering jobs include:

Infographic showing various Remote Chaos Engineering job openings in the United States as of August 2026, with employment types broken down into 78% Full Time, 11% Temporary, and 11% Contract. Highlights an 100% Remote job distribution, with an average salary of $194,709 per year, or $93.6 per hour.

Site Reliability Engineer/ Chaos Engineer Remote to start

Georgia IT, Inc.

Remote

$58.25 - $77.50/hr

Full-time

Re-posted 5 days ago


Job description

Job Summary:
Georgia IT, Inc. is looking for a Site Reliability Engineer/Chaos Engineer to join their SRE team. The role involves leading technical efforts to enhance the reliability and stability of mission-critical systems through Chaos Engineering practices and collaboration with various teams.
Responsibilities:
• Create operational tooling for monitoring self-healing infrastructures and chaos testing.
• Design and create controlled chaos in production systems.
• Create Python and Terraform scripts.
• Work across teams identify and fix issues that affect systems reliability and performance.
• Guide and design architectural decisions and direct solutions that will enhance our client s product reliability.
• Dive into system and latent reliability issues service performance and capacity modeling of distributed systems at scale.
• Partner with development team to identify anti patterns and optimization strategies create fallback options and help develop self-healing capabilities across the enterprise in a sustainable manner.
Qualifications:
Required:
• Candidate will be part of the SRE team and lead technical role to determine Reliability Chaos Engineering needs of mission critical systems and business processes.
• Candidate will assess high level architecture and design issues relating to platform enterprise software interactions with other systems.
• Application development infrastructure database and middleware teams to ensure stability and reliability of the system.
• Chaos Engineering will proactively detect issues within the applications platform network and databases in a controlled way using Chaos tools like Chaos Monkey Gremlin Simian Army.
• Candidate should have familiarity with Internet protocols such as HTTP DNS TCP and UDP and Linux development environment and well versed with DevOps.
• Candidate will identify anti patterns optimization and support development of self-healing capabilities.
• Create operational tooling for monitoring self-healing infrastructures and chaos testing.
• Design and create controlled chaos in production systems.
• Create Python and Terraform scripts.
• Work across teams identify and fix issues that affect systems reliability and performance.
• Guide and design architectural decisions and direct solutions that will enhance our client's product reliability.
• Dive into system and latent reliability issues service performance and capacity modeling of distributed systems at scale.
• Partner with development team to identify anti patterns and optimization strategies create fallback options and help develop self-healing capabilities across the enterprise in a sustainable manner.
Company:
Georgia IT, Inc. provides IT Consulting for a wide range of IT services and custom build turn-key enterprise solutions. Founded in 2007, the company is headquartered in Alpharetta, USA, with a team of 51-200 employees. The company is currently Growth Stage.

Georgia IT logo

About Georgia IT

Sourced by ZipRecruiter

A PROFESSIONAL SERVICES ORGANIZATION WITH A VISION OF DELIVERING SIMPLE AFFORDABLE, SUSTAINABLE SOLUTIONS FOR COMPLEX PROBLEMS WITH INTEGRITY. OUR GOAL IS TO ACHIEVE ALL THIS IN A COLLABRATIVE APPROACH WITH ALL PARTIES INVOLVED IN DELIVERING SOLUTIONS/PRODUCTS.

Industry

It services

Company size

51 - 200 Employees

Headquarters location

Alpharetta, GA, US

Year founded

2007

Social media