1

Chaos Engineering Jobs in Georgia (NOW HIRING)

SRE Lead/ Architect

Atlanta, GA · On-site

$54.75 - $72.75/hr

Mandatory skills are Observability, Resiliency, Chaos engineering, strong python, and Dynatrace As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and ...

Experience with chaos engineering tools (Gremlin, Chaos Monkey) * Background in product-facing services with high traffic scale * Understand how to use incident management platforms. This includes ...

Senior Software Engineer - SRE

Atlanta, GA · On-site

$54.75 - $72.75/hr

Experience with chaos engineering tools (Gremlin, Chaos Monkey) * Background in product-facing services with high traffic scale * Understand how to use incident management platforms. This includes ...

Senior AI Platform SWE - PlayerZero

Atlanta, GA · On-site

$117K - $155K/yr

... chaos engineering, live migrations, and deep protections for customer data. • Partner closely with ML researchers and product leads to translate novel ideas into hardened infrastructure.

Software Engineer

Atlanta, GA · On-site

$116K - $174K/yr

Experience with chaos engineering tools (Gremlin, Chaos Monkey) * Background in product-facing services with high traffic scale * Understand how to use incident management platforms. This includes ...

Software Engineer

Atlanta, GA · On-site

$116K - $174K/yr

Experience with chaos engineering tools (Gremlin, Chaos Monkey) * Background in product-facing services with high traffic scale * Understand how to use incident management platforms. This includes ...

Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents. * For example: in your first few ...

Site Reliability Engineer

Alpharetta, GA

$55.75 - $74/hr

Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents. * For example: in your first few ...

Site Reliability Engineer

Alpharetta, GA · On-site

$55.75 - $74/hr

Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents. * For example: in your first few ...

Senior Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

Experience with chaos engineering or resiliency testing * Experience with high-volume, high-availability transactional systems * Experience with AI-assisted observability or operational automation

Senior Site Reliability Engineer

Atlanta, GA

$54.75 - $72.75/hr

Experience with chaos engineering or resiliency testing * Experience with high-volume, high-availability transactional systems * Experience with AI-assisted observability or operational automation

Senior Cloud Devops Engineer

Atlanta, GA · On-site

$125K - $160K/yr

... practices and chaos engineering background. * Open source contributions relevant to the DevOps or cloud ecosystem. * Active certifications including CKA, CKAD, AWS Certified DevOps Engineer ...

next page

Showing results 1-20

Chaos Engineering information

What is chaos engineering?

A Chaos Engineering job involves proactively identifying weaknesses in complex systems by intentionally injecting failures and observing how they respond. Professionals in this role design and execute controlled experiments to improve system resilience, ensuring that services remain reliable under unexpected conditions. They work closely with development, operations, and security teams to enhance fault tolerance and incident response strategies.

What are some typical challenges a chaos engineer faces, and how do they overcome them?

Chaos Engineers often face the challenge of designing effective experiments that simulate real-world failures without disrupting production systems. Balancing the need to discover vulnerabilities with maintaining uptime requires careful planning, communication, and coordination with development and operations teams. They address these challenges by thoroughly testing in controlled environments, documenting procedures, and establishing clear rollback strategies. Continuous learning and cross-functional collaboration are also key to staying ahead of new complexities in evolving systems.

What are the key skills and qualifications needed to thrive in the chaos engineering position, and why are they important?

To thrive in Chaos Engineering, a strong background in software engineering, distributed systems, and reliability testing is essential, often supported by a degree in computer science or a related field. Familiarity with chaos engineering tools like Gremlin or Chaos Monkey and experience with cloud platforms, container orchestration, and monitoring systems are highly valued. Excellent problem-solving abilities, communication skills, and a mindset oriented toward experimentation help engineers collaborate effectively and analyze complex failure modes. These skills are crucial for proactively identifying system weaknesses and ensuring the resilience of large-scale technology infrastructures.

Is chaos engineering still relevant?

Chaos engineering is a valuable practice for proactively identifying system vulnerabilities by intentionally introducing failures. It remains relevant in modern DevOps and cloud environments to improve system resilience and reliability, often utilizing tools like Chaos Monkey and Gremlin. As systems grow more complex, the need for chaos engineering skills continues to increase for engineers focused on fault tolerance and system stability.

What does a chaos engineer do?

A chaos engineer designs and executes experiments to intentionally disrupt systems in order to identify vulnerabilities and improve resilience. They use tools like chaos engineering frameworks to simulate failures and ensure systems can withstand unexpected issues, often working closely with development and operations teams. Strong knowledge of distributed systems, scripting, and monitoring is essential for this role.

What are the most commonly searched types of Chaos Engineering jobs in Georgia?

The most popular types of Chaos Engineering jobs in Georgia are:

What are popular job titles related to Chaos Engineering jobs in Georgia?

For Chaos Engineering jobs in Georgia, the most frequently searched job titles are:

What job categories do people searching Chaos Engineering jobs in Georgia look for?

The top searched job categories for Chaos Engineering jobs in Georgia are:

What cities in Georgia are hiring for Chaos Engineering jobs?

Cities in Georgia with the most Chaos Engineering job openings:

Infographic showing various Chaos Engineering job openings in Georgia as of August 2026, with employment types broken down into 53% Full Time, and 47% Contract. Highlights an 100% In-person job distribution.

SRE Lead/ Architect

Vish Consulting IT

Atlanta, GA • On-site

$54.75 - $72.75/hr

Contractor

Re-posted 20 days ago


Job description

Job Title: SRE Lead/Architect

Location: Atlanta, GA  - Hybrid (Thur to next wed (Alternate weeks))

Contract Role

Role Summary: Mandatory skills are Observability, Resiliency, Chaos engineering, strong python, and Dynatrace

As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems and practices that ensure the reliability, scalability, performance, and efficiency of our critical services. Moving beyond day-to-day operations, you will focus on the strategic architectural direction of SRE function, defining standards, blueprints, and frameworks that enable development teams and fellow SRE operations team to build and operate highly resilient systems. Leverage deep expertise in software engineering, distributed systems, cloud infrastructure, and SRE principles to influence technology choices, establish best practices, and foster a proactive culture of reliability across the organization and much beyond observability pillar.

Key Responsibilities:

  1. Reliability Strategy & Design:
    • Architect and design highly available, scalable, secure, and cost-effective infrastructure and application patterns on AWS
    • Define and evangelize SRE best practices, standards, and blueprints for service design, deployment, monitoring, and operational readiness across the engineering organization
    • Review current observability implementation to identify gaps and define steps to reach next level maturity of observability setup  to provide deep insights into system health and behaviour
    • With overall maturity lead the definition and implementation strategy for Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets for critical services
  1. Platform Architecture & Automation:
    • Design solutions to systematically reduce operational toil through automation and improved system design
    • Evaluate current SRE tools and automation frameworks (e.g., CI/CD pipelines, Infrastructure as Code modules, automated incident remediation, chaos engineering platforms) and suggest enhancement that will help overall enhancement of capability
    • Evaluate, prototype, and recommend new technologies, tools, and methodologies to enhance system reliability, developer productivity, and operational efficiency
  1. Technical Leadership & Consultation:
    • Act as a senior technical advisor and subject matter expert on reliability, scalability, and performance for development and platform teams
    • Provide architectural guidance during the design phase of new services and features to ensure reliability principles are embedded early (shift-left)
    • Mentor and coach other SREs and engineers, fostering technical excellence and adherence to SRE principles
    • Lead architectural reviews and production readiness assessments for critical systems
  1. Resilience:
    • Lead blameless postmortems for significant incidents, ensuring root causes are identified and systemic architectural improvements are prioritized and implemented
    • Architect and advocate for resilience patterns (e.g., circuit breaking, rate limiting, graceful degradation, chaos engineering) within applications and infrastructure

Required Qualifications:

  • Proven experience in an architectural role, designing solutions for reliability, scalability, and performance
  • Deep understanding and practical application of SRE principles (SLIs/SLOs, error budgets, toil reduction, automation, incident management, postmortems)
  • Expertise in cloud computing platforms (e.g., AWS) including infrastructure, networking, and security services
  • Strong experience with containerization and orchestration technologies (Kubernetes, Docker, serverless computing)
  • Solid experience designing and implementing observability solutions (e.g., Dynatrace, Prometheus, Grafana, ELK/EFK Stack, Jaeger, OpenTelemetry)
  • Strong programming/scripting skills (e.g., Python, Go, Bash) for automation and tool development
  • Excellent analytical, problem-solving, and strategic thinking skills.
  • Strong communication, collaboration, and leadership skills with the ability to influence technical direction across teams

Preferred Qualifications:

  • Experience designing and implementing chaos engineering practices and platforms

Thanks & Regards,

Vivek Sharma 

Account Manager

Cell: (904) 481-0481

Fax: (619)-333-1294

Email: vivek@vishusa.com

Vish Consulting Services, Inc

www.vishusa.com