Data Scientist, AI/ML
$220K - $290K/yr
As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world's largest organizations where high availability is non-negotiable. About the Role of the Data ...
$220K - $290K/yr
As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world's largest organizations where high availability is non-negotiable. About the Role of the Data ...
$220K - $290K/yr
As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world's largest organizations where high availability is non-negotiable. About the Role of the Data ...
Durham, NC · On-site
$120 - $150/hr
Integrates Site Reliability Engineering (SRE) practices (Observability and Chaos) with DevOps processes and delivery pipelines to stop bad code from reaching production. Ensures business‑critical ...
Durham, NC · On-site
$120 - $150/hr
Integrates Site Reliability Engineering (SRE) practices (Observability and Chaos) with DevOps processes and delivery pipelines to stop bad code from reaching production. Ensures business‑critical ...
$44 - $58.50/hr
Establish a continuous chaos engineering and resilience testing practice through fault injection, game days, and controlled failure experiments. * Mentor senior and mid-level engineers while raising ...
$44 - $58.50/hr
Establish a continuous chaos engineering and resilience testing practice through fault injection, game days, and controlled failure experiments. * Mentor senior and mid-level engineers while raising ...
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
El Segundo, CA · On-site
$54K - $73K/yr
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
El Segundo, CA · On-site
$54K - $73K/yr
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
You will partner closely with CHAOS engineering teams to ensure customer priorities inform internal roadmaps, and coordinate with Service Growth Leads (Army, Air Force, Navy, etc.) where relevant to ...
You will partner closely with CHAOS engineering teams to ensure customer priorities inform internal roadmaps, and coordinate with Service Growth Leads (Army, Air Force, Navy, etc.) where relevant to ...
$54K - $73K/yr
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
$54K - $73K/yr
You will lead basic system troubleshooting with confidence and support CHAOS engineering teams in the diagnosis and resolution of hardware, software, networking, and system integration issues.
Washington, DC · On-site
$137K - $177K/yr
You will partner closely with CHAOS engineering teams to ensure customer priorities inform internal roadmaps, and coordinate with Service Growth Leads (Army, Air Force, Navy, etc.) where relevant to ...
Washington, DC · On-site
$137K - $177K/yr
You will partner closely with CHAOS engineering teams to ensure customer priorities inform internal roadmaps, and coordinate with Service Growth Leads (Army, Air Force, Navy, etc.) where relevant to ...
... chaos engineering. Responsibilities Own the enterprise Performance Engineering charter as a horizontal shared service; define engagement models, intake/prioritization, and performance-readiness ...
Quick apply
... chaos engineering. Responsibilities Own the enterprise Performance Engineering charter as a horizontal shared service; define engagement models, intake/prioritization, and performance-readiness ...
You will partner closely with CHAOS engineering teams to ensure customer priorities inform internal roadmaps, and coordinate with Service Growth Leads (Army, Air Force, Navy, etc.) where relevant to ...
You will partner closely with CHAOS engineering teams to ensure customer priorities inform internal roadmaps, and coordinate with Service Growth Leads (Army, Air Force, Navy, etc.) where relevant to ...
Austin, TX · On-site +1
$208K - $244K/yr
... design, chaos engineering, and recovery runbook ownership • Build and operate observability and resilience tooling to ensure infrastructure state is fully instrumented, drift is detected ...
Austin, TX · On-site +1
$208K - $244K/yr
... design, chaos engineering, and recovery runbook ownership • Build and operate observability and resilience tooling to ensure infrastructure state is fully instrumented, drift is detected ...
You will partner closely with CHAOS engineering teams to ensure customer priorities inform internal roadmaps, and coordinate with Service Growth Leads (Army, Air Force, Navy, etc.) where relevant to ...
You will partner closely with CHAOS engineering teams to ensure customer priorities inform internal roadmaps, and coordinate with Service Growth Leads (Army, Air Force, Navy, etc.) where relevant to ...
... Chaos Engineering experiments to validate system stability under failure conditions. · Provide performance tuning recommendations to development and infrastructure teams. · Prepare detailed ...
Quick apply
... Chaos Engineering experiments to validate system stability under failure conditions. · Provide performance tuning recommendations to development and infrastructure teams. · Prepare detailed ...
Austin, TX · On-site
$130 - $160/hr
Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing * Participate in game day exercises and disaster recovery simulations; create ...
Austin, TX · On-site
$130 - $160/hr
Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing * Participate in game day exercises and disaster recovery simulations; create ...
Implement reliability engineering practices such as chaos engineering and proactive risk mitigation to reduce operational risk and improve resilience. Job responsibilities * Vision, strategy, and ...
Implement reliability engineering practices such as chaos engineering and proactive risk mitigation to reduce operational risk and improve resilience. Job responsibilities * Vision, strategy, and ...
Wood Ridge, NJ · On-site
$58 - $77.25/hr
Apply chaos engineering principles to test system resilience and recovery. * Participate in on-call rotations and incident response, ensuring quick restoration of service. Required Skills ...
Wood Ridge, NJ · On-site
$58 - $77.25/hr
Apply chaos engineering principles to test system resilience and recovery. * Participate in on-call rotations and incident response, ensuring quick restoration of service. Required Skills ...
Implement reliability engineering practices such as chaos engineering and proactive risk mitigation to reduce operational risk and improve resilience. Job responsibilities * Vision, strategy, and ...
Implement reliability engineering practices such as chaos engineering and proactive risk mitigation to reduce operational risk and improve resilience. Job responsibilities * Vision, strategy, and ...
Implement reliability engineering practices such as chaos engineering and proactive risk mitigation to reduce operational risk and improve resilience. Job responsibilities * Vision, strategy, and ...
Implement reliability engineering practices such as chaos engineering and proactive risk mitigation to reduce operational risk and improve resilience. Job responsibilities * Vision, strategy, and ...
$46.5K - $58.1K
1% of jobs
$58.1K - $69.7K
1% of jobs
$69.7K - $81.3K
2% of jobs
$81.3K - $92.9K
5% of jobs
$92.9K - $104.5K
7% of jobs
$114.1K is the 25th percentile. Wages below this are outliers.
$104.5K - $116K
10% of jobs
$116K - $127.6K
7% of jobs
$127.6K - $139.2K
6% of jobs
$139.2K - $150.8K
4% of jobs
$150.8K - $162.4K
3% of jobs
The median wage is $162.9K / yr.
$162.4K - $174K
52% of jobs
$46.5K
$146.9K
$174K
A Chaos Engineering job involves proactively identifying weaknesses in complex systems by intentionally injecting failures and observing how they respond. Professionals in this role design and execute controlled experiments to improve system resilience, ensuring that services remain reliable under unexpected conditions. They work closely with development, operations, and security teams to enhance fault tolerance and incident response strategies.
Chaos Engineers often face the challenge of designing effective experiments that simulate real-world failures without disrupting production systems. Balancing the need to discover vulnerabilities with maintaining uptime requires careful planning, communication, and coordination with development and operations teams. They address these challenges by thoroughly testing in controlled environments, documenting procedures, and establishing clear rollback strategies. Continuous learning and cross-functional collaboration are also key to staying ahead of new complexities in evolving systems.
To thrive in Chaos Engineering, a strong background in software engineering, distributed systems, and reliability testing is essential, often supported by a degree in computer science or a related field. Familiarity with chaos engineering tools like Gremlin or Chaos Monkey and experience with cloud platforms, container orchestration, and monitoring systems are highly valued. Excellent problem-solving abilities, communication skills, and a mindset oriented toward experimentation help engineers collaborate effectively and analyze complex failure modes. These skills are crucial for proactively identifying system weaknesses and ensuring the resilience of large-scale technology infrastructures.
Cities with the most Chaos Engineering job openings:
The most popular types of Chaos Engineering jobs are:
States with the most job openings for Chaos Engineering jobs include:
The top searched job categories for Chaos Engineering jobs are:

Today's complex, fast-paced systems have become a minefield of reliability risks, any of which could cause an outage that costs millions and destroys customer confidence. That's why high-availability teams use Gremlin to find and fix reliability risks before they become incidents.
Gremlin Reliability Platform helps software teams proactively monitor and test their systems for common reliability risks, build and enforce reliability standards, and automate their reliability practices organization-wide. As the industry leader in Chaos Engineering and reliability testing, we work with hundreds of the world's largest organizations where high availability is non-negotiable.
About the Role of the Data Scientist, AI/MLAs a Data Scientist, AI/ML at Gremlin, you will have the opportunity to improve the reliability of the internet at large by turning millions of chaos engineering experiments into automated failure analysis and remediation. You will be able to leverage your applied machine learning experience to inform product direction as well as solve complex technical problems that directly impact our customers (which range from the Fortune 500 to smaller organizations). You will work closely with a small, talented engineering team focused on quality, delivery, and predictability with an emphasis on providing our customers a great user experience.
In this role, you'll get to:*The role does not offer sponsorship employment benefits.
**If you don't think you meet all of the criteria above but still are interested in the job, please apply. Nobody checks every box, we're looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.
Compensation
We expect the salary range for this role to be $220,000 - $290,000. We recognize that salary varies from person to person depending on level of experience and we welcome direct conversations about it. The final offer will vary based on assessment of a candidate's skills and ability and our budget and market data.
Gremlin offers competitive total compensation packages including 401k Matching, Equity and other benefits such as flexible time off and paid company holidays.
About Gremlin:Gremlin is a team of industry veterans and people eager to learn from one another. We set the standard for reliability and equip leading organizations with the mindset and expertise needed to drive reliability improvements that move the world forward. We're backed by top-tier investors Index Ventures, Amplify Partners, and Redpoint Ventures. Our customers love us, and we're thrilled to be a partner in their success.
What Do We Care About:We Care about our People
People are our critical differentiators. The company strives to treat our people with respect, empathy, and dignity. We expect that our people will treat each other similarly. In both cases, we will assume good intent. All are welcome at Gremlin. We know our differences make us stronger and that our best ideas and contributions can come from anyone at any level.
We Care about Collaboration
Gremlin is strongest when we come together as one team with shared goals. Be the glue, not the glitter. But as a remote company, teamwork and collaboration won't happen by accident. We approach every challenge as a shared challenge. We rely on each other for diverse perspectives and creative ideas. We celebrate our wins as a team.
We Care about Results
Be high productivity, low drama. Results matter. To keep our pace, everyone owns the outcomes of their actions and takes action when needed. We reward speed over perfection. We empower each other to iterate and experiment. You are welcome at Gremlin for who you are. The more voices and ideas we have represented in our business, the more we will all flourish, contribute, and build a more reliable internet.
Gremlin is a place where everyone can grow and is encouraged. However you identify and whatever background you bring with you, please apply if this sounds like a role that would make you excited to come into work everyday. It's in our differences that we will find the power to keep building a more reliable internet by building and designing tools used by the best companies in the world.
Visit our website to learn more - https://www.gremlin.com/about
Sourced by ZipRecruiter
Software development
11 - 50 Employees
Bodega Bay, CA, US
2016