2

Remote Chaos Engineering Jobs in Florida (NOW HIRING)

Senior DevOps Engineer, Infrastructure & Reliability

Tampa, FL ยท Remote

$122K - $157K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments ... SLOs, error budgets, chaos testing). All Remote Hires will be required to travel to Orlando ...

Senior DevOps Engineer, Infrastructure & Reliability

Miami, FL ยท Remote

$124K - $159K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Partner with engineering teams to eliminate friction in CI/CD, deployments, and cloud environments ... SLOs, error budgets, chaos testing). All Remote Hires will be required to travel to Orlando ...

The Senior Director of Engineering will lead a large global engineering organization responsible ... They should bring structure without bureaucracy, urgency without chaos, and innovation without ...

Remote Chaos Engineering information

What does a remote chaos engineering do?

A remote chaos engineering professional designs and executes experiments to intentionally disrupt systems in order to identify vulnerabilities and improve resilience. They use tools like Chaos Monkey or Gremlin and often work with cloud environments, monitoring system behavior to ensure reliability and fault tolerance. Strong scripting skills and understanding of distributed systems are essential for this role.

What is remote chaos engineering?

Remote Chaos Engineering is the practice of testing distributed systems' resilience by intentionally introducing failures and disruptions in remote or cloud environments. The goal is to identify weaknesses and improve system reliability by simulating real-world incidents, such as network outages or server crashes, in a controlled manner. This approach helps teams understand how their applications behave under stress and develop strategies to mitigate future incidents. Remote Chaos Engineering is particularly valuable for organizations leveraging cloud infrastructure and remote services, ensuring robust performance even under unexpected conditions.

What are some common challenges faced by professionals working in remote chaos engineering roles?

Professionals in remote chaos engineering often encounter challenges such as coordinating experiments across distributed teams, ensuring clear communication about system vulnerabilities, and managing the complexity of large-scale systems without direct, on-site access. Establishing robust monitoring and rollback procedures is essential to minimize risk during remote testing. Additionally, building trust with development and operations teams is key, as chaos engineering often involves intentionally introducing failures to improve system resilience.

Which remote chaos engineering jobs can be done remotely?

Remote chaos engineering jobs are commonly available in roles such as Site Reliability Engineer, DevOps Engineer, or SRE, which often involve designing and testing system resilience using tools like Chaos Monkey or Gremlin. These positions typically require strong scripting skills and familiarity with cloud platforms, and they can often be performed entirely remotely depending on the company's policies.

What are the key skills and qualifications needed to thrive as a remote chaos engineer?

To thrive as a Remote Chaos Engineer, you need a strong background in software engineering, systems architecture, and site reliability, often supported by a degree in computer science or a related field. Familiarity with chaos engineering platforms (such as Gremlin or Chaos Monkey), cloud environments (AWS, Azure, GCP), and automation tools is typically required. Strong problem-solving abilities, clear communication, and a collaborative mindset help you effectively identify weaknesses and drive reliability improvements across distributed teams. These skills are crucial for proactively uncovering system vulnerabilities, ensuring system resilience, and maintaining high availability in complex, remote-first infrastructures.

What is the difference between Remote Chaos Engineering vs Remote Site Reliability Engineer?

AspectRemote Chaos EngineeringRemote Site Reliability Engineer
Primary FocusDesigning and executing chaos experiments to improve system resilienceEnsuring system reliability, availability, and performance through monitoring and automation
Skills & CertificationsKnowledge of chaos engineering tools, scripting, cloud platformsMonitoring tools, scripting, cloud infrastructure, SRE certifications
Work EnvironmentCollaborates with development and operations teams, often in DevOps cultureWorks closely with engineering teams to maintain system health and SLAs

While both roles focus on system stability, Remote Chaos Engineering specializes in testing system resilience through chaos experiments, whereas Remote Site Reliability Engineers focus on maintaining overall system reliability and performance. Both roles require scripting skills and cloud knowledge, but their core objectives differ: one proactively tests, the other maintains system health.

What are the most commonly searched types of Chaos Engineering jobs in Florida?

The most popular types of Chaos Engineering jobs in Florida are:

What are popular job titles related to Remote Chaos Engineering jobs in Florida?

For Remote Chaos Engineering jobs in Florida, the most frequently searched job titles are:

What job categories do people searching Remote Chaos Engineering jobs in Florida look for?

The top searched job categories for Remote Chaos Engineering jobs in Florida are:

What cities in Florida are hiring for Remote Chaos Engineering jobs?

Cities in Florida with the most Remote Chaos Engineering job openings:

Infographic showing various Remote Chaos Engineering job openings in Florida as of August 2026, with employment types broken down into 90% Full Time, 5% Part Time, 4% Contract, and 1% Nights. Highlights an 86% Physical, 4% Hybrid, and 10% Remote job distribution.

Site Reliability Engineer

Bayswater Consulting Group

Miami, FL โ€ข Remote

$130K - $170K/yr

Full-time

This job post hasย expired today.ย Applications are no longer accepted.


Job description

Site Reliability Engineer (SRE)

About the Role

We are working with a client seeking an experienced Site Reliability Engineer for a full-time, direct hire opportunity. This is a fully remote position open to candidates across the United States.

Responsibilities

  • Own reliability, availability, and performance of production systems and services
  • Define and track service level objectives (SLOs), service level indicators (SLIs), and error budgets
  • Build and maintain observability infrastructure — logging, metrics, tracing, and alerting
  • Develop automation to eliminate toil and improve operational efficiency across the engineering organization
  • Lead incident response — on-call rotation, root cause analysis, and post-mortem documentation
  • Collaborate with development teams to build reliability into services from the ground up
  • Design and implement disaster recovery and business continuity capabilities
  • Manage and improve CI/CD pipelines and deployment infrastructure
  • Champion engineering best practices around reliability, scalability, and security

Requirements

  • 4+ years of experience in a site reliability, platform engineering, or senior DevOps role
  • Strong software engineering foundation — proficiency in Python, Go, or similar
  • Deep experience with cloud infrastructure — AWS, Azure, or GCP
  • Hands-on experience with Kubernetes and container orchestration at scale
  • Proficiency with observability and monitoring tools — Datadog, Prometheus, Grafana, PagerDuty, or similar
  • Experience with infrastructure as code — Terraform, Pulumi, or similar
  • Strong understanding of distributed systems, networking, and system design
  • Proven incident management and on-call experience
  • Excellent communication skills — able to work across engineering, product, and leadership

Nice to Have

  • Experience implementing SRE practices from scratch in a growing organization
  • Familiarity with chaos engineering tools such as Gremlin or LitmusChaos
  • Experience with service mesh technologies such as Istio or Linkerd
  • Background in large-scale, high-availability production environments
  • Knowledge of FinOps and cloud cost optimization
  • Experience mentoring junior engineers on reliability practices

Compensation

$130,000 — $170,000 base salary depending on experience. Full benefits package available.

Location

Fully remote — United States.