2

Remote Chaos Engineering Jobs in Arizona (NOW HIRING)

next page

Showing results 1-20

Remote Chaos Engineering information

What does a remote chaos engineering do?

A remote chaos engineering professional designs and executes experiments to intentionally disrupt systems in order to identify vulnerabilities and improve resilience. They use tools like Chaos Monkey or Gremlin and often work with cloud environments, monitoring system behavior to ensure reliability and fault tolerance. Strong scripting skills and understanding of distributed systems are essential for this role.

What is remote chaos engineering?

Remote Chaos Engineering is the practice of testing distributed systems' resilience by intentionally introducing failures and disruptions in remote or cloud environments. The goal is to identify weaknesses and improve system reliability by simulating real-world incidents, such as network outages or server crashes, in a controlled manner. This approach helps teams understand how their applications behave under stress and develop strategies to mitigate future incidents. Remote Chaos Engineering is particularly valuable for organizations leveraging cloud infrastructure and remote services, ensuring robust performance even under unexpected conditions.

What are some common challenges faced by professionals working in remote chaos engineering roles?

Professionals in remote chaos engineering often encounter challenges such as coordinating experiments across distributed teams, ensuring clear communication about system vulnerabilities, and managing the complexity of large-scale systems without direct, on-site access. Establishing robust monitoring and rollback procedures is essential to minimize risk during remote testing. Additionally, building trust with development and operations teams is key, as chaos engineering often involves intentionally introducing failures to improve system resilience.

Which remote chaos engineering jobs can be done remotely?

Remote chaos engineering jobs are commonly available in roles such as Site Reliability Engineer, DevOps Engineer, or SRE, which often involve designing and testing system resilience using tools like Chaos Monkey or Gremlin. These positions typically require strong scripting skills and familiarity with cloud platforms, and they can often be performed entirely remotely depending on the company's policies.

What are the key skills and qualifications needed to thrive as a remote chaos engineer?

To thrive as a Remote Chaos Engineer, you need a strong background in software engineering, systems architecture, and site reliability, often supported by a degree in computer science or a related field. Familiarity with chaos engineering platforms (such as Gremlin or Chaos Monkey), cloud environments (AWS, Azure, GCP), and automation tools is typically required. Strong problem-solving abilities, clear communication, and a collaborative mindset help you effectively identify weaknesses and drive reliability improvements across distributed teams. These skills are crucial for proactively uncovering system vulnerabilities, ensuring system resilience, and maintaining high availability in complex, remote-first infrastructures.

What is the difference between Remote Chaos Engineering vs Remote Site Reliability Engineer?

AspectRemote Chaos EngineeringRemote Site Reliability Engineer
Primary FocusDesigning and executing chaos experiments to improve system resilienceEnsuring system reliability, availability, and performance through monitoring and automation
Skills & CertificationsKnowledge of chaos engineering tools, scripting, cloud platformsMonitoring tools, scripting, cloud infrastructure, SRE certifications
Work EnvironmentCollaborates with development and operations teams, often in DevOps cultureWorks closely with engineering teams to maintain system health and SLAs

While both roles focus on system stability, Remote Chaos Engineering specializes in testing system resilience through chaos experiments, whereas Remote Site Reliability Engineers focus on maintaining overall system reliability and performance. Both roles require scripting skills and cloud knowledge, but their core objectives differ: one proactively tests, the other maintains system health.

What are the most commonly searched types of Chaos Engineering jobs in Arizona? The most popular types of Chaos Engineering jobs in Arizona are:
What are popular job titles related to Remote Chaos Engineering jobs in Arizona? For Remote Chaos Engineering jobs in Arizona, the most frequently searched job titles are:
What cities in Arizona are hiring for Remote Chaos Engineering jobs? Cities in Arizona with the most Remote Chaos Engineering job openings:
Infographic showing various Remote Chaos Engineering job openings in Arizona as of June 2026, with employment types broken down into 1% As Needed, 11% Full Time, 61% Part Time, 26% Contract, and 1% Nights. Highlights an 89% Physical, 3% Hybrid, and 8% Remote job distribution.

REMOTE - AI Engineering Manager (Databricks)

State Farm

Phoenix, AZ • On-site, Remote

Full-time

Medical, Dental, Vision, Retirement

Re-posted 22 days ago


State Farm rating

7.3

Company rating: 7.3 out of 10

Based on 1,549 frontline employees who took The Breakroom Quiz

226th of 304 rated insurance


Job description

Overview

Being good neighbors – helping people, investing in our communities, and making the world a better place – is who we are at State Farm. It is at the core of how we operate and the reason for our success. Come join a #1 team and do some good!

REMOTE:Qualified candidates living outside a 50-mile radius of a hub location listed below can plan to work remote. 

HYBRID: Qualified candidates living within a 50-mile radius of a hub location listed below should plan to spend time working from home and some time working in the office as part of our hybrid work environment.
HUB LOCATIONS: Bloomington, IL; Dunwoody, GA; Richardson, TX; or Tempe, AZ 

SPONSORSHIP:  Applicants for this position are required to be eligible to lawfully work in the U.S. immediately; employer will not sponsor applicants for U.S. work authorization (e.g. H-1B visa) for this opportunity.


Responsibilities

Lead a team of 3-5 embedded agentic engineers who work inside Product Oriented Delivery pods, helping engineers, analysts, and product owners ship twice as fast with twice the quality through agentic workflows. You'll build the agentic harness tuned to our existing infrastructure, write code alongside your team, run the developer community, and ensure a steady stream of innovation projects make it to production on Databricks

What You'll Own

  • Team Leadership (30%) – Hire, manage, and grow 3-5 embedded engineers. 1:1s, career development, removing blockers. You code 30-40% of the time.
  • Agentic Harness (25%) – Build the agentic harness tuned to our infrastructure – hooks, connectors, and integrations with Claude for code quality checks, artifact generation, and next-best-action guidance through the SDLC.
  • Databricks Solutions (20%) – Production templates for Medallion pipelines, Unity Catalog governance, MLflow, PySpark. Not PoCs – real systems with monitoring and runbooks.
  • Developer Community (15%) – Demos, office hours, pattern libraries. Measure adoption and impact.
  • Production Readiness (10%) – Quality gates, automated checks, incident coaching, ET handoff.

What Success Looks Like

  • Team hired, embedded in pods, and shipping agentic infrastructure
  • Developer adoption >70% – teams actively using agentic tools in daily work
  • Projects consistently shipping to production with proper monitoring and handoff
  • Cycle time cut in half. Incidents down 50%. Measured, not estimated.
  • Developer community thriving with demos, office hours, and a living pattern library
  • You've developed a successor on your team

Stack

Databricks (Delta Lake, Unity Catalog, MLflow, PySpark) / Anthropic Claude API / Python, SQL / AWS (Lambda, ECS, RDS) / DataDog, Prometheus, Grafana / Git / CLI tools, VS Code extensions, CI/CD hooks

Culture

  • Kindness is the standard. In the code, the docs, the commit messages.
  • Objective over agreement. Debate hard, commit fully.
  • Agentic-first. If an LLM can augment it, we do it.
  • Production or nothing. We ship.
  • Blameless accountability. RCAs that teach, not punish.
  • Player-coach model. Leaders build alongside the team.
  • 2x2 is the mission. Measured, not aspirational.

Qualifications

Must-Haves

  • Leadership: 2+ years managing engineering teams. Experience with embedded/distributed teams. Coaching mindset.
  • Production (Non-Negotiable): Shipped end-to-end systems to production with real users, SLAs, and on-call. Not PoCs. 1+ year operating production ML/AI systems.
  • Databricks: 2+ years production Databricks (Delta Lake, Unity Catalog, MLflow, PySpark). Medallion architecture. Cost optimization.
  • Agentic/LLMs: Anthropic Claude production experience (structured outputs, tool use, multi-turn). Built developer tools with LLMs. Evaluation frameworks. Prompt engineering at scale.
  • Engineering: Production-grade Python. SQL expertise. CI/CD. API design. Code quality obsession. Threat modeling. Chaos engineering.
  • Communication: Technical teaching. Influence without authority. Clear written communication. Stakeholder management.

Nice-to-Haves

  • Developer platforms or CLI tools. DORA/SPACE metrics. Open-source contributions. Event streaming. Privacy regulations. Insurance/fintech/regulated industries. Lightweight UI skills (Streamlit, Gradio, FastAPI + HTMX).

Our Benefits

Because work-life balance is a priority at State Farm, compensation is based on our standard 38:45-hour work week!

  • Potential starting salary range: $151,000 - $247,000. Starting salary will be based on skills, background, and experience. High end of the range limited to applicants with significant relevant experience and living in high cost of labor location like CA/NY/NJ/Etc. 
  • Potential yearly incentive pay up to 18% of base salary

At State Farm, we offer more than just a paycheck. Check out our suite of benefits designed to give you the flexibility you need to take care of you and your family!

  • Get Paid! On top of our competitive pay, you are eligible for an annual raise and bonus.
  • Stay Well! Focus on you and your family’s health with our robust health and wellbeing programs. State Farm pays most of your healthcare premium, and we offer multiple healthcare plan options, including a high deductible plan. All medical plans provide 100% coverage for in-network preventative care, AND you and your family have access to vision, dental, telemedicine, 24/7 mental health professionals, and much more!
  • Develop and Grow! Take advantage of educational benefits like industry leading training programs, top-notch tuition assistance programs, employee resource groups, and mentoring.
  • Plan Ahead! Plan for those big moments in life with benefits like fertility/IVF/adoption assistance, college coaching, national discount programs, interactive monthly financial workshops, free financial coaching, and more. You can also start a savings account or consider financing through our State Farm Federal Credit Union!
  • Take a Little “You” Time! You will have access to our generous time off policies designed so you can plan around holidays, family events, volunteering, or just to take a relaxing day off. With the opportunity to initially earn up to 20 days annually plus parental leave, paid holidays, celebration day, life leave (40 hours/year), bereavement leave, and community service/education support days, there will be plenty of time for you!
  • Give Back! We offer several ways to give back through our Matching Gift Program, Good Neighbor Grant Program, and the Employee Assistance Fund.
  • Finish Strong! Plan for retirement using free financial advisors and a 401(k) plan with company contributions of up to 7% of your salary.

Visit our State Farm Careers page for more information on our benefits, locations, and the hiring process of joining the State Farm team!

Qualifications:

Must-Haves

  • Leadership: 2+ years managing engineering teams. Experience with embedded/distributed teams. Coaching mindset.
  • Production (Non-Negotiable): Shipped end-to-end systems to production with real users, SLAs, and on-call. Not PoCs. 1+ year operating production ML/AI systems.
  • Databricks: 2+ years production Databricks (Delta Lake, Unity Catalog, MLflow, PySpark). Medallion architecture. Cost optimization.
  • Agentic/LLMs: Anthropic Claude production experience (structured outputs, tool use, multi-turn). Built developer tools with LLMs. Evaluation frameworks. Prompt engineering at scale.
  • Engineering: Production-grade Python. SQL expertise. CI/CD. API design. Code quality obsession. Threat modeling. Chaos engineering.
  • Communication: Technical teaching. Influence without authority. Clear written communication. Stakeholder management.

Nice-to-Haves

  • Developer platforms or CLI tools. DORA/SPACE metrics. Open-source contributions. Event streaming. Privacy regulations. Insurance/fintech/regulated industries. Lightweight UI skills (Streamlit, Gradio, FastAPI + HTMX).
Education:UNAVAILABLEEmployment Type: FULL_TIME

What State Farm employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom