2

Remote Chaos Engineering Jobs in Georgia (NOW HIRING)

Senior Cloud Devops Engineer

Atlanta, GA · On-site +1

$125K - $160K/yr

Strong Terraform skills with production-grade IaC, modules, and remote state management. * Hands-on ... and chaos engineering background. * Open source contributions relevant to the DevOps or cloud ...

Client Director (GA, NC, MA, NY, CA, Remote) Aimpoint Digital is a fast-growing, independently ... Being a "builder" and appreciating the chaos while enjoying the rewards is a critical quality ...

Client Director (GA, NC, MA, NY, CA, Remote) Aimpoint Digital is a fast-growing, independently ... Being a "builder" and appreciating the chaos while enjoying the rewards is a critical quality ...

Remote Chaos Engineering information

What is remote chaos engineering?

Remote Chaos Engineering is the practice of testing distributed systems' resilience by intentionally introducing failures and disruptions in remote or cloud environments. The goal is to identify weaknesses and improve system reliability by simulating real-world incidents, such as network outages or server crashes, in a controlled manner. This approach helps teams understand how their applications behave under stress and develop strategies to mitigate future incidents. Remote Chaos Engineering is particularly valuable for organizations leveraging cloud infrastructure and remote services, ensuring robust performance even under unexpected conditions.

What are the key skills and qualifications needed to thrive as a remote chaos engineer?

To thrive as a Remote Chaos Engineer, you need a strong background in software engineering, systems architecture, and site reliability, often supported by a degree in computer science or a related field. Familiarity with chaos engineering platforms (such as Gremlin or Chaos Monkey), cloud environments (AWS, Azure, GCP), and automation tools is typically required. Strong problem-solving abilities, clear communication, and a collaborative mindset help you effectively identify weaknesses and drive reliability improvements across distributed teams. These skills are crucial for proactively uncovering system vulnerabilities, ensuring system resilience, and maintaining high availability in complex, remote-first infrastructures.

What are some common challenges faced by professionals working in remote chaos engineering roles?

Professionals in remote chaos engineering often encounter challenges such as coordinating experiments across distributed teams, ensuring clear communication about system vulnerabilities, and managing the complexity of large-scale systems without direct, on-site access. Establishing robust monitoring and rollback procedures is essential to minimize risk during remote testing. Additionally, building trust with development and operations teams is key, as chaos engineering often involves intentionally introducing failures to improve system resilience.

What is the difference between Remote Chaos Engineering vs Remote Site Reliability Engineer?

AspectRemote Chaos EngineeringRemote Site Reliability Engineer
Primary FocusDesigning and executing chaos experiments to improve system resilienceEnsuring system reliability, availability, and performance through monitoring and automation
Skills & CertificationsKnowledge of chaos engineering tools, scripting, cloud platformsMonitoring tools, scripting, cloud infrastructure, SRE certifications
Work EnvironmentCollaborates with development and operations teams, often in DevOps cultureWorks closely with engineering teams to maintain system health and SLAs

While both roles focus on system stability, Remote Chaos Engineering specializes in testing system resilience through chaos experiments, whereas Remote Site Reliability Engineers focus on maintaining overall system reliability and performance. Both roles require scripting skills and cloud knowledge, but their core objectives differ: one proactively tests, the other maintains system health.

What job categories do people searching Remote Chaos Engineering jobs in Georgia look for?

The top searched job categories for Remote Chaos Engineering jobs in Georgia are:

What cities in Georgia are hiring for Remote Chaos Engineering jobs?

Cities in Georgia with the most Remote Chaos Engineering job openings:

Senior Cloud Devops Engineer

Atlanta, GA • On-site, Remote

Verifone
Insurance Carriers • 5 - 10K employees

$125K - $160K/yr

Full-time

Re-posted yesterday


Job description

Why Verifone

For more than 30 years, Verifone has established a remarkable record of leadership in the electronic payment technology industry. Verifone has one of the leading electronic payment solutions brands and is one of the largest providers of electronic payment systems worldwide.

Verifone has a diverse, dynamic and fast paced work environment in which employees are focused on results and have opportunities to excel. We take pride in the fact that we work with leading retailers, merchants, banks, and third-party partners to invent and deliver innovative payments solution around the world. We strive for excellence in our products and services and are obsessed with customer happiness. Across the globe, Verifone employees are leading the payments industry through experience, innovation, and an ambitious spirit. Whether it's developing the next platform of secure payment systems or searching for new ways to bring electronic payments to new markets, the team at Verifone is dedicated to the success of our customers, partners and investors. It is this passion for innovation that drives each one of our employees for personal and professional success.

Verifone is proudly an in-office work culture as we see immense benefits to career development and business results from our colleagues being physically co-located.

What's exciting about the role

The Senior DevOps Cloud Engineer is a hands-on technical role at the core of Verifone's cloud infrastructure function. You will design, build, and maintain scalable AWS infrastructure, drive automation through Infrastructure as Code, and champion Kubernetes-based container orchestration at scale - all while partnering closely with development teams to enable rapid, safe software delivery.This is a role for a cloud-native engineer who brings deep AWS and Kubernetes expertise, a strong automation mindset, and the drive to continuously improve reliability, security, and developer experience across a global engineering organization.

Skills and Experience we desire

AWS & Cloud Infrastructure

  • Design and manage scalable AWS infrastructure across EC2, EKS, RDS, S3, IAM, VPC, CloudWatch, and Lambda.
  • Own AWS account governance - including cost optimization, security posture, and multi-account strategies.
  • Implement and maintain cloud networking, load balancing, and DNS configurations.
  • Enforce AWS security best practices including IAM policies, secrets management, and compliance controls.

Kubernetes & Container Orchestration

  • Lead the design, deployment, and operations of Kubernetes clusters on Amazon EKS.
  • Manage workload scheduling, autoscaling (HPA/VPA/KEDA), resource quotas, and namespace strategies.
  • Troubleshoot cluster and pod-level issues including networking (CNI), ingress, and storage.

Infrastructure as Code

  • Build and maintain modular Terraform codebases for AWS and EKS provisioning.
  • Manage Terraform state, modules, and workspaces across a team environment.
  • Enforce IaC code review standards and GitOps workflows.

CI/CD & Automation

  • Design and maintain CI/CD pipelines using Bitbucket Actions, Jenkins, and ArgoCD.
  • Automate operational tasks and reduce toil using Python scripting.
  • Build tooling to improve developer experience and deployment velocity.

Monitoring & Reliability

  • Implement and maintain observability stacks including New Relic, CloudWatch, and Elastic.
  • Participate in on-call rotations, incident response, and post-mortems.
  • Drive SLO/SLA definitions and ongoing reliability improvements.
Qualifications
  • 5+ years of experience in a DevOps, SRE, or Cloud Engineering role.
  • Deep AWS expertise across compute, networking, storage, and security services.
  • Expert-level Kubernetes knowledge - cluster administration, workload management, networking, and troubleshooting.
  • Strong Terraform skills with production-grade IaC, modules, and remote state management.
  • Hands-on Amazon EKS experience including managed node groups, Fargate, add-ons, and upgrades.
  • Solid Python scripting for automation and operational tooling.
  • Strong networking fundamentals - VPCs, subnets, security groups, VPNs, and DNS.

Preferred Experience
  • Helm, Istio/Linkerd, ArgoCD/Flux, or HashiCorp Vault experience.
  • SRE practices and chaos engineering background.
  • Open source contributions relevant to the DevOps or cloud ecosystem.
  • Active certifications including CKA, CKAD, AWS Certified DevOps Engineer - Professional, or AWS Certified Solutions Architect.
Our commitment

Verifone is committed to creating a diverse environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, or veteran status. Verifone is also committed to compliance with all fair employment practices regarding citizenship and immigration status.