1

Chaos Engineering Jobs in Texas (NOW HIRING)

Sr Site Reliability Engineer

Austin, TX · On-site

$56.50 - $75/hr

Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing * Participate in game day exercises and disaster recovery simulations; create ...

Sr Site Reliability Engineer

Austin, TX · On-site

$56.50 - $75/hr

Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing * Participate in game day exercises and disaster recovery simulations; create ...

Sr Site Reliability Engineer

Austin, TX

$56.50 - $75/hr

Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing * Participate in game day exercises and disaster recovery simulations; create ...

New

Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing * Participate in game day exercises and disaster recovery simulations; create ...

Software Performance engineer

Dallas, TX · On-site

$138K/yr

Any experience with Chaos Engineering is a plus * We know the confidence gap and imposter syndrome can get in the way of meeting spectacular candidates. Please don't hesitate to apply. Manikanth ...

Senior .NET Engineer

Austin, TX · On-site

$121K - $160K/yr

Collaborate Site Reliability Engineering (SRE) on Chaos Engineering and Observability for your project. Qualifications: * Education: * Bachelor's or Master's degree in computer science, Engineering ...

Senior .NET Engineer

Austin, TX · On-site

$121K - $160K/yr

Collaborate Site Reliability Engineering (SRE) on Chaos Engineering and Observability for your project. Qualifications: * Education: * Bachelor's or Master's degree in computer science, Engineering ...

Senior Site Reliability Engineer

Dallas, TX · On-site

$56.50 - $75/hr

Automation & Resilience Engineering: Experience with automation and resilience practices such as Python-based automation, RPA platforms (e.g., Blue Prism, UiPath), chaos engineering, and failure ...

next page

Showing results 1-20

Chaos Engineering information

What is chaos engineering?

A Chaos Engineering job involves proactively identifying weaknesses in complex systems by intentionally injecting failures and observing how they respond. Professionals in this role design and execute controlled experiments to improve system resilience, ensuring that services remain reliable under unexpected conditions. They work closely with development, operations, and security teams to enhance fault tolerance and incident response strategies.

What are some typical challenges a chaos engineer faces, and how do they overcome them?

Chaos Engineers often face the challenge of designing effective experiments that simulate real-world failures without disrupting production systems. Balancing the need to discover vulnerabilities with maintaining uptime requires careful planning, communication, and coordination with development and operations teams. They address these challenges by thoroughly testing in controlled environments, documenting procedures, and establishing clear rollback strategies. Continuous learning and cross-functional collaboration are also key to staying ahead of new complexities in evolving systems.

What are the key skills and qualifications needed to thrive in the chaos engineering position, and why are they important?

To thrive in Chaos Engineering, a strong background in software engineering, distributed systems, and reliability testing is essential, often supported by a degree in computer science or a related field. Familiarity with chaos engineering tools like Gremlin or Chaos Monkey and experience with cloud platforms, container orchestration, and monitoring systems are highly valued. Excellent problem-solving abilities, communication skills, and a mindset oriented toward experimentation help engineers collaborate effectively and analyze complex failure modes. These skills are crucial for proactively identifying system weaknesses and ensuring the resilience of large-scale technology infrastructures.

Is chaos engineering still relevant?

Chaos engineering is a valuable practice for proactively identifying system vulnerabilities by intentionally introducing failures. It remains relevant in modern DevOps and cloud environments to improve system resilience and reliability, often utilizing tools like Chaos Monkey and Gremlin. As systems grow more complex, the need for chaos engineering skills continues to increase for engineers focused on fault tolerance and system stability.

What does a chaos engineer do?

A chaos engineer designs and executes experiments to intentionally disrupt systems in order to identify vulnerabilities and improve resilience. They use tools like chaos engineering frameworks to simulate failures and ensure systems can withstand unexpected issues, often working closely with development and operations teams. Strong knowledge of distributed systems, scripting, and monitoring is essential for this role.

What are the most commonly searched types of Chaos Engineering jobs in Texas?

The most popular types of Chaos Engineering jobs in Texas are:

What cities in Texas are hiring for Chaos Engineering jobs?

Cities in Texas with the most Chaos Engineering job openings:

Infographic showing various Chaos Engineering job openings in Texas as of August 2026, with employment types broken down into 59% Full Time, and 41% Contract. Highlights an 100% In-person job distribution.

Sr Site Reliability Engineer

Austin, TX • On-site

$56.50 - $75/hr

Full-time

Medical, Dental, Vision, Retirement

Re-posted 6 days ago


Job description

Recognized as the No. 1 site trusted by real estate professionals, Realtor.com® has been at the forefront of online real estate for over 25 years, connecting buyers, sellers, and renters with trusted insights and expert guidance to find their perfect home. Through its robust suite of tools, Realtor.com® not only makes a significant impact on the real estate industry at large, but for consumers, navigating the biggest purchase they will make in their life, by providing a user experience that is easy to use, easy to understand, and most of all, easy to make decisions.
Join us on our mission to empower more people to find their way home by breaking barriers to entry, making the right connections, and building confidence through expert guidance.
We are seeking a Senior Site Reliability Engineer to join our newly formed Operations Excellence organization, reporting to the Director, Operations Excellence. This role will contribute to the reliability, observability, and operational excellence of our platform infrastructure serving millions of users. As a Senior SRE, you will be a strong technical contributor who implements best practices, solves complex problems, and enables our 600+ engineers to deliver exceptional customer experiences. You will work on critical platform systems including EKS infrastructure, Skyway (CI/CD), Frontdoor (Tyk API Gateway), Pantheon (Apollo GraphQL Federation), and our observability stack, while contributing to chaos engineering practices and cost optimization initiatives with measurable ROI.
What You'll Do:
Platform Reliability & Infrastructure
  • Implement and maintain highly available AWS infrastructure including EKS clusters, Fargate (ECS), and multi-region architectures
  • Support reliability of critical services: Skyway (CI/CD), Frontdoor (Tyk), Pantheon (Apollo GraphQL), and supporting infrastructure
  • Monitor SLIs, SLOs, and error budgets for Tier 1/2/3 systems; participate in architectural reviews for reliability and cost-efficiency
  • Implement reliability patterns including circuit breakers, graceful degradation, and automated failover

Observability & Cost Optimization
  • Implement observability solutions using NewRelic for APM, distributed tracing, metrics, and logging for rapid troubleshooting
  • Build dashboards and alerts that reduce MTTD and MTTR; contribute to observability standards across teams
  • Identify infrastructure cost optimization opportunities and implement FinOps practices including rightsizing and resource lifecycle management
  • Support cost-conscious architecture decisions and CI/CD spend optimization (CircleCI, Argo CD)

Chaos Engineering & Incident Response
  • Execute chaos engineering experiments to identify system weaknesses; contribute to frameworks for safe production testing
  • Participate in game day exercises and disaster recovery simulations; create runbooks and automation for resilience
  • Participate in on-call rotation for critical systems; conduct post-incident reviews and implement improvements
  • Support incident response processes and contribute to System Health Scorecard

Technical Contribution
  • Contribute as a strong technical individual contributor to the Operations Excellence team
  • Collaborate with Platform Engineering, Quality Engineering, and product teams on reliability initiatives
  • Support security initiatives including AWS Secrets Manager migration and compliance requirements (SOC 2, PCI, GDPR)
  • Contribute to Developer Experience metrics and platform adoption goals
  • May provide technical guidance to junior team members

What You'll Bring:
  • 5+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering with demonstrated success improving system reliability
  • Bachelor's degree or equivalent experience
  • 3+ years hands-on experience with AWS (EKS, EC2, RDS, S3, CloudWatch, IAM) and Kubernetes including cluster management
  • Proficient programming skills (Python, Go, or Java) with infrastructure automation and Infrastructure as Code experience (Terraform, CloudFormation)
  • Production experience with observability tools (NewRelic, Datadog, Prometheus, Grafana, Splunk) and distributed systems
  • Experience with CI/CD platforms and GitOps workflows (CircleCI, Argo CD, Jenkins); on-call rotation and incident response
  • Preferred: Exposure to chaos engineering tools, API Gateway technologies (Tyk/Kong), GraphQL federation (Apollo), cost optimization initiatives, FinOps principles

Technical Skills
  • Cloud & Infrastructure: AWS (EKS, Fargate, Lambda, VPC, Route53, CloudFront), Kubernetes, Docker, Istio Service Mesh
  • CI/CD & GitOps: Argo CD, CircleCI, Jenkins, GitHub Actions
  • Observability: NewRelic - APM, distributed tracing, metrics & logging; Splunk - logging
  • IaC & Automation: Terraform, CloudFormation, Helm, Kustomize, Python/Go/Bash
  • Platform Services: Tyk Gateway, Apollo GraphQL, AWS Secrets Manager, Vault
  • Incident Management: OpsGenie, PagerDuty, ServiceNow

Professional Qualities
  • Strong communication skills with ability to explain technical concepts to diverse audiences
  • Collaborative approach working across engineering, product, and business teams
  • Self-motivated with ability to solve complex problems within established practices and policies
  • Data-driven decision making with customer-centric approach and empathy for developer experience

How We Work:
We balance creativity and innovation on a foundation of in-person collaboration. For most roles, our employees work three or more days in our offices, where they have the opportunity to collaborate in-person, adding richness to our culture and knitting us closer together.
How We Reward You:
Realtor.com is committed to investing in the health and wellbeing of our employees and their families. Our benefits programs include, but are not limited to:
  • Inclusive and Competitive medical, Rx, dental, and vision coverage
  • Family forming benefits
  • 13 Paid Holidays
  • Flexible Time Off
  • 8 hours of paid Volunteer Time off
  • Immediate eligibility into Company 401(k) plan with 3.5% company match
  • Tuition Reimbursement program for degreed and non-degreed programs
  • 1:1 personalized Financial Planning Sessions
  • Student Debt Retirement Savings Match program
  • Free snacks and refreshments in each office location

Do the best work of your life at Realtor.com®
Here, you'll partner with a diverse team of experts as you use leading-edge tech to empower everyone to meet a crucial goal: finding their way home. And you'll find your way home too. At Realtor.com®, you'll bring your full self to work as you innovate with speed, serve our consumers, and champion your teammates. In return, we'll provide you with a warm, welcoming, and inclusive culture; intellectual challenges; and the development opportunities you need to grow.
Diversity is important to us, therefore, Realtor.com® is an Equal Opportunity Employer regardless of age, color, national origin, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, marital status, status as a disabled veteran and/or veteran of the Vietnam Era or any other characteristic protected by federal, state or local law. In addition, Realtor.com® will provide reasonable accommodations for otherwise qualified disabled individuals.