1

Chaos Engineering Jobs in Florida (NOW HIRING)

Lead Java Developer - Senior Vice President

Tampa, FL

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Experience with chaos engineering, fault injection, and production readiness reviews as part of a resilience testing practice. What we offer Joining Citi means taking ownership of engineering that ...

Enterprise Observability Architect

Tampa, FL · On-site

$62.75 - $81/hr

  • Medical

  • Life

  • Retirement

  • PTO

Lead enterprise assessments, failure-mode analysis, chaos engineering practices, and post-incident improvement cycles * Translate telemetry insights into business-level narratives that inform risk ...

next page

Showing results 1-20

Chaos Engineering information

See Florida salary details

$34.7K

$109.8K

$130K

How much do chaos engineering jobs pay per year?

As of Aug 17, 2026, the average yearly pay for chaos engineering in Florida is $109,753.00, according to ZipRecruiter salary data. Most workers in this role earn between $87,100.00 and $129,300.00 per year, depending on experience, location, and employer.

What is chaos engineering?

A Chaos Engineering job involves proactively identifying weaknesses in complex systems by intentionally injecting failures and observing how they respond. Professionals in this role design and execute controlled experiments to improve system resilience, ensuring that services remain reliable under unexpected conditions. They work closely with development, operations, and security teams to enhance fault tolerance and incident response strategies.

What are some typical challenges a chaos engineer faces, and how do they overcome them?

Chaos Engineers often face the challenge of designing effective experiments that simulate real-world failures without disrupting production systems. Balancing the need to discover vulnerabilities with maintaining uptime requires careful planning, communication, and coordination with development and operations teams. They address these challenges by thoroughly testing in controlled environments, documenting procedures, and establishing clear rollback strategies. Continuous learning and cross-functional collaboration are also key to staying ahead of new complexities in evolving systems.

What are the key skills and qualifications needed to thrive in the chaos engineering position, and why are they important?

To thrive in Chaos Engineering, a strong background in software engineering, distributed systems, and reliability testing is essential, often supported by a degree in computer science or a related field. Familiarity with chaos engineering tools like Gremlin or Chaos Monkey and experience with cloud platforms, container orchestration, and monitoring systems are highly valued. Excellent problem-solving abilities, communication skills, and a mindset oriented toward experimentation help engineers collaborate effectively and analyze complex failure modes. These skills are crucial for proactively identifying system weaknesses and ensuring the resilience of large-scale technology infrastructures.

Is chaos engineering still relevant?

Chaos engineering is a valuable practice for proactively identifying system vulnerabilities by intentionally introducing failures. It remains relevant in modern DevOps and cloud environments to improve system resilience and reliability, often utilizing tools like Chaos Monkey and Gremlin. As systems grow more complex, the need for chaos engineering skills continues to increase for engineers focused on fault tolerance and system stability.

What does a chaos engineer do?

A chaos engineer designs and executes experiments to intentionally disrupt systems in order to identify vulnerabilities and improve resilience. They use tools like chaos engineering frameworks to simulate failures and ensure systems can withstand unexpected issues, often working closely with development and operations teams. Strong knowledge of distributed systems, scripting, and monitoring is essential for this role.

What are the most commonly searched types of Chaos Engineering jobs in Florida?

The most popular types of Chaos Engineering jobs in Florida are:

What job categories do people searching Chaos Engineering jobs in Florida look for?

The top searched job categories for Chaos Engineering jobs in Florida are:

Infographic showing various Chaos Engineering job openings in Florida as of August 2026, with employment types broken down into 58% Full Time, and 42% Contract. Highlights an 100% In-person job distribution, with an average salary of $109,753 per year, or $52.8 per hour.

Information Technology_USA - USA_Engineer

Real Soft, Inc.

Jacksonville, FL • On-site

$52.75 - $70.25/hr

Contractor

Re-posted 16 days ago


Job description

**Please strictly adhere to the following resume naming convention:
ALL CAPS, NO SPACES BETWEEN UNDERSCORES
PTN_US_GBAMSREQID_CandidateBeelineID
Example: PTN_US_9999999_SKIPJOHNSON0413
: -
MSP Owner: Michelle Lee
Location: Long Lake, WA
Duration: 6 months
skill id: 10878828
Role Description
• Experience supporting large-scale distributed systems (200+ microservices).
• Strong hands-on experience with FluxCD for GitOps-based continuous delivery.
• Experience implementing and managing:
o GitOps deployment patterns
o Automated Kubernetes application reconciliation
o Multi-environment deployment strategies
o Progressive delivery and automated rollbacks
o Configuration and secret management
o Helm-based application deployments through FluxCD
• Expertise integrating FluxCD with Enterprise GitHub, Jenkins, Azure Container Registry (ACR), and Kubernetes clusters.
• Experience managing platform releases for 200+ microservices using GitOps principles.
• Ability to troubleshoot deployment drift, reconciliation failures, and cluster synchronization issues.
• Strong understanding of:
o SLI/SLO/SLA management
o Incident Management
o Problem Management
o Capacity Planning
o Reliability Engineering
o Root Cause Analysis (RCA)
o Chaos Engineering principles
Required Skills
Site Reliability Engineering
• Reliability Engineering
• SLI/SLO/SLA management
• Incident Management
• Problem Management
• Capacity Planning
• Root Cause Analysis (RCA)
• Chaos Engineering principles
Kubernetes & GitOps
• Kubernetes
• FluxCD
• GitOps deployment patterns
• Automated Kubernetes application reconciliation
• Multi-environment deployment strategies
• Progressive delivery
• Automated rollbacks
• Helm-based application deployments through FluxCD
• Configuration and secret management
Platform Operations
• Experience supporting large-scale distributed systems (200+ microservices)
• Experience managing platform releases for 200+ microservices using GitOps principles
• Troubleshooting deployment drift
• Reconciliation failures
• Cluster synchronization issues
DevOps & Automation
• CI/CD automation
• Terraform
• GitHub
• Jenkins
• Azure Container Registry (ACR)
Application & Messaging Technologies
• Java ecosystem
• Kafka
• ActiveMQ
• Event-driven architectures
• Modern UI applications
Observability & Database Operations
• Observability
• Database operations
Cloud-Native Platforms
• Cloud-native environments
Candidate Profile
• We are seeking a highly skilled Senior Platform Site Reliability Engineer (SRE) to support and operate a large-scale distributed platform consisting of 200+ Java-based microservices, event-driven architectures, modern UI applications, and enterprise-grade databases.
• The candidate will drive platform reliability, scalability, observability, automation, and operational excellence across cloud-native environments.
• The ideal engineer will have strong expertise in Kubernetes, FluxCD, Terraform, Java ecosystem, Kafka/ActiveMQ messaging, CI/CD automation, observability, and database operations, with a mindset focused on reliability engineering, automation, and production support.
Skills: Digital : Kafka ~ Digital : Kubernetes ~ Digital: Terraform ~ Java Scheduler Tools - Quartz| Flux ~ ActiveMQ
Experience Required: 10 & Above