2

Remote Chaos Engineering Jobs in Virginia (NOW HIRING)

Familiarity with chaos engineering practices and tooling. * Experience with data pipeline ... This is a remote position. While performing the duties of this job, the employee regularly works in ...

Familiarity with chaos engineering practices and tooling. * Experience with data pipeline ... This is a remote position. While performing the duties of this job, the employee regularly works in ...

Remote Chaos Engineering information

What does a remote chaos engineering do?

A remote chaos engineering professional designs and executes experiments to intentionally disrupt systems in order to identify vulnerabilities and improve resilience. They use tools like Chaos Monkey or Gremlin and often work with cloud environments, monitoring system behavior to ensure reliability and fault tolerance. Strong scripting skills and understanding of distributed systems are essential for this role.

What is remote chaos engineering?

Remote Chaos Engineering is the practice of testing distributed systems' resilience by intentionally introducing failures and disruptions in remote or cloud environments. The goal is to identify weaknesses and improve system reliability by simulating real-world incidents, such as network outages or server crashes, in a controlled manner. This approach helps teams understand how their applications behave under stress and develop strategies to mitigate future incidents. Remote Chaos Engineering is particularly valuable for organizations leveraging cloud infrastructure and remote services, ensuring robust performance even under unexpected conditions.

What are some common challenges faced by professionals working in remote chaos engineering roles?

Professionals in remote chaos engineering often encounter challenges such as coordinating experiments across distributed teams, ensuring clear communication about system vulnerabilities, and managing the complexity of large-scale systems without direct, on-site access. Establishing robust monitoring and rollback procedures is essential to minimize risk during remote testing. Additionally, building trust with development and operations teams is key, as chaos engineering often involves intentionally introducing failures to improve system resilience.

Which remote chaos engineering jobs can be done remotely?

Remote chaos engineering jobs are commonly available in roles such as Site Reliability Engineer, DevOps Engineer, or SRE, which often involve designing and testing system resilience using tools like Chaos Monkey or Gremlin. These positions typically require strong scripting skills and familiarity with cloud platforms, and they can often be performed entirely remotely depending on the company's policies.

What are the key skills and qualifications needed to thrive as a remote chaos engineer?

To thrive as a Remote Chaos Engineer, you need a strong background in software engineering, systems architecture, and site reliability, often supported by a degree in computer science or a related field. Familiarity with chaos engineering platforms (such as Gremlin or Chaos Monkey), cloud environments (AWS, Azure, GCP), and automation tools is typically required. Strong problem-solving abilities, clear communication, and a collaborative mindset help you effectively identify weaknesses and drive reliability improvements across distributed teams. These skills are crucial for proactively uncovering system vulnerabilities, ensuring system resilience, and maintaining high availability in complex, remote-first infrastructures.

What is the difference between Remote Chaos Engineering vs Remote Site Reliability Engineer?

AspectRemote Chaos EngineeringRemote Site Reliability Engineer
Primary FocusDesigning and executing chaos experiments to improve system resilienceEnsuring system reliability, availability, and performance through monitoring and automation
Skills & CertificationsKnowledge of chaos engineering tools, scripting, cloud platformsMonitoring tools, scripting, cloud infrastructure, SRE certifications
Work EnvironmentCollaborates with development and operations teams, often in DevOps cultureWorks closely with engineering teams to maintain system health and SLAs

While both roles focus on system stability, Remote Chaos Engineering specializes in testing system resilience through chaos experiments, whereas Remote Site Reliability Engineers focus on maintaining overall system reliability and performance. Both roles require scripting skills and cloud knowledge, but their core objectives differ: one proactively tests, the other maintains system health.

What are the most commonly searched types of Chaos Engineering jobs in Virginia?

The most popular types of Chaos Engineering jobs in Virginia are:

What are popular job titles related to Remote Chaos Engineering jobs in Virginia?

For Remote Chaos Engineering jobs in Virginia, the most frequently searched job titles are:

What job categories do people searching Remote Chaos Engineering jobs in Virginia look for?

The top searched job categories for Remote Chaos Engineering jobs in Virginia are:

What cities in Virginia are hiring for Remote Chaos Engineering jobs?

Cities in Virginia with the most Remote Chaos Engineering job openings:

Infographic showing various Remote Chaos Engineering job openings in Virginia as of August 2026, with employment types broken down into 90% Full Time, 5% Part Time, 4% Contract, and 1% Nights. Highlights an 86% Physical, 4% Hybrid, and 10% Remote job distribution.

Sr Site Reliability Engineer

Commence

Virginia Beach, VA • On-site, Remote

$145K - $175K/yr

Full-time

Posted 9 days ago


Job description

Description

At Commence, we're the start of a new age of data-centric transformation, elevating health outcomes and powering better, more efficient process to program and patient health. We combine quality data-driven solutions that fuel answers, technology that advances performance, and clinical expertise that builds trust to create a more efficient path to quality care. 


With human-centered, healthcare-relevant, and value-based solutions, we create new possibilities with data. We provide proof beyond the concept and performance beyond the scope with a focus on efficiencies that transform the lives of those we serve. With a culture driven by purpose, straightforward communication and clinical domain expertise, Commence cuts straight to better care.

Requirements

As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and operational health of our mission-critical healthcare data platform. You will bridge the gap between engineering and operations-embedding reliability as a first-class concern from architecture through deployment. This role is built for someone who thrives when systems are under pressure and who treats an outage as a problem to be engineered away permanently, not just survived. 

  • Design, implement, and own observability infrastructure including metrics, logging, tracing, and alerting across distributed systems.
  • Define and enforce SLOs, SLIs, and error budgets in partnership with product and engineering teams.
  • Lead incident response: triage, coordinate remediation, conduct blameless post-mortems, and drive systemic fixes.
  • Build and maintain CI/CD pipelines that support rapid, safe delivery of changes to production.
  • Collaborate with engineering teams on infrastructure changes; able to read, modify, and contribute to existing infrastructure-as-code (Terraform or CloudFormation).
  • Design and operate highly available, fault-tolerant systems-including auto-scaling, failover, and disaster recovery strategies.
  • Reduce operational toil through automation; eliminate manual processes before they become habits.
  • Collaborate with software engineers to establish reliability-first design patterns and review architectures for operational risk.
  • Manage Kubernetes or container orchestration environments at scale.
  • Ensure systems meet compliance and security requirements, particularly those applicable to healthcare data (HIPAA, SOC 2).
  • Provide technical mentorship and guidance to engineers across the organization on reliability practices.
  • Participate in on-call rotation with a commitment to continuously reducing the need for it.

Qualifications

  • 7+ years of experience in SRE, platform engineering, or DevOps roles.
  • Exceptional problem-solving under pressure-demonstrated track record of diagnosing complex, high-stakes system failures and building durable solutions.
  • Deep hands-on experience with AWS services including EC2, EKS/ECS, Lambda, RDS, S3, CloudWatch, and related tooling.
  • Familiarity with infrastructure-as-code (Terraform or CloudFormation)-able to contribute to existing configurations.
  • Experience designing and operating distributed systems with strict availability and latency requirements.
  • Proficiency in at least one scripting or systems language (Python, Go, Bash, or similar) for automation and tooling.
  • Experience with container orchestration (Kubernetes, ECS) in production environments.
  • Expertise in observability tooling (OpenSearch, Prometheus/Grafana, or equivalent).
  • Hands-on experience with CI/CD platforms (GitHub Actions, Jenkins, CircleCI, or similar).
  • Proven ability to define and operationalize SLOs and error budgets.
  • Experience with relational and NoSQL databases-performance tuning, replication, and backup strategies.
  • Strong working knowledge of networking fundamentals: DNS, load balancing, VPCs, TLS.
  • Excellent communication skills-able to translate technical risk into business impact for non-engineering stakeholders.

Additional Requirements

  • AWS Certifications (Solutions Architect, DevOps Engineer, or SysOps Administrator).
  • Experience in healthcare technology or other regulated industries (HIPAA, SOC 2, FedRAMP).
  • Familiarity with chaos engineering practices and tooling.
  • Experience with data pipeline reliability (ETL/ELT workflows, streaming systems).
  • Exposure to AI/ML infrastructure and the reliability challenges unique to model serving.
  • Familiarity with additional cloud platforms (Azure, Google Cloud).
  • Contributions to open-source reliability or infrastructure tooling.

*Commence' headquarters are in Virginia Beach, VA, however we are open to remote candidates in the following states:   AZ, AR, CO, DE, FL, GA, IL, IN, KS, KY, MA, MD, MI, MS, MO, MT, NC, NE, NV, NY, OH, OK, PA, SC, TN, TX, VA, DC, WI, and WV* 

Work Environment/Physical Demands

The work environment and physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.


This is a remote position. While performing the duties of this job, the employee regularly works in a climate-controlled environment. Candidates must be able to sit, read, work on a computer, and watch a computer screen for extended periods of time. Occasionally required to stand, walk, use hands and fingers, kneel or crouch.


Commence is an equal employment opportunity employer. All personnel processes are merit-based and applied without discrimination on the basis of race, color, religion, sex, sexual orientation, gender identity, marital status, age, disability, national or ethnic origin, military and veteran status or any other characteristic protected by applicable law. 


Commence.AI is committed to providing equal employment opportunities to all applicants, including individuals with disabilities. If you require a reasonable accommodation to participate in the application process due to a disability, please contact Human Resources at (757) 306-4920 or hr@commence.ai. Please note that unless you are requesting an accommodation, all applications must be submitted through our online application system.