1

Director Chaos Engineering Jobs in Ridgewood, NJ

Senior Production Engineer

New York, NY · On-site

$139K - $204K/yr

We are hiring a Senior Production Engineer to take direct, hands-on ownership of critical tooling ... Familiarity with DR/BCP, service tiering, capacity planning, or chaos engineering. * Background ...

Senior Production Engineer

Livingston, NJ · On-site

$139K - $204K/yr

We are hiring a Senior Production Engineer to take direct, hands-on ownership of critical tooling ... Familiarity with DR/BCP, service tiering, capacity planning, or chaos engineering. * Background ...

Senior Production Engineer

New York, NY · On-site

$139K - $204K/yr

We are hiring a Senior Production Engineer to take direct, hands-on ownership of critical tooling ... Familiarity with DR/BCP, service tiering, capacity planning, or chaos engineering. * Background ...

Senior Production Engineer

Livingston, NJ · On-site

$139K - $204K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

We are hiring a Senior Production Engineer to take direct, hands-on ownership of critical tooling ... Familiarity with DR/BCP, service tiering, capacity planning, or chaos engineering. * Background ...

Chaos Labs - Full Stack Engineer (AI)

New York, NY · On-site +1

$160K - $210K/yr

  • Medical

  • Dental

  • Vision

  • PTO

Comfortable working in fast-paced, self-directed environments. Preferred Qualifications: * Based in ... Familiarity with prompt engineering, tool use orchestration (MCP), and AI agent design.

next page

Showing results 1-20

Director Chaos Engineering information

See Ridgewood, NJ salary details

$73.9K

$197K

$257K

How much do director chaos engineering jobs pay per year?

As of Aug 16, 2026, the average yearly pay for director chaos engineering in Ridgewood, NJ is $197,004.00, according to ZipRecruiter salary data. Most workers in this role earn between $143,200.00 and $256,000.00 per year, depending on experience, location, and employer.

What are some common challenges faced by a director of chaos engineering when implementing chaos experiments at scale?

A Director of Chaos Engineering often encounters challenges such as gaining buy-in from stakeholders who may be unfamiliar or skeptical about deliberately introducing failures. Coordinating chaos experiments across multiple teams and complex distributed systems requires careful planning to ensure safety and minimize unintended disruptions. Balancing the need for rigorous experimentation with ongoing business priorities and service reliability can also be difficult. Building a culture that values resilience and learning from failure is essential for long-term success in this role.

What are the key skills and qualifications needed to thrive as a director of chaos engineering?

To thrive as a Director of Chaos Engineering, you need deep expertise in distributed systems, site reliability engineering, and a strong background in computer science or a related field. Familiarity with chaos engineering platforms (such as Gremlin or Chaos Monkey), cloud infrastructure (AWS, Azure, GCP), and relevant certifications like AWS Certified Solutions Architect are typically required. Exceptional leadership, problem-solving, and communication skills help drive cross-team collaboration and foster a culture of resilience. These skills and qualities are essential to proactively identify system weaknesses, minimize downtime, and ensure business continuity in complex technical environments.

What is a director of chaos engineering?

A Director of Chaos Engineering is a senior technology leader responsible for overseeing the development and implementation of chaos engineering practices within an organization. Their primary goal is to improve system resilience by proactively testing how systems respond to failures and unexpected disruptions. This role involves designing and leading experiments that intentionally introduce faults to identify vulnerabilities, guiding teams in building more robust systems, and fostering a culture of reliability. They often collaborate with engineering, operations, and security teams to ensure best practices are followed across the organization.

What is the difference between Director Chaos Engineering vs Site Reliability Engineer?

AspectDirector Chaos EngineeringSite Reliability Engineer
CredentialsAdvanced certifications in chaos engineering, cloud platforms, and leadershipCertifications in SRE, cloud, and DevOps tools
Work EnvironmentLeadership role overseeing chaos testing strategies across teamsOperational role managing system reliability and automation
Industry UsageUsed in organizations focusing on resilience and fault toleranceCommon in tech companies maintaining scalable, reliable systems
Search & ComparisonOften compared for strategic impact and leadership scopeCompared for technical expertise and system management

The Director Chaos Engineering focuses on leading chaos testing initiatives and shaping resilience strategies, while the Site Reliability Engineer handles the day-to-day reliability and automation of systems. Both roles are vital in ensuring system robustness but differ in scope and responsibilities.

What job categories do people searching Director Chaos Engineering jobs in Ridgewood, NJ look for?

The top searched job categories for Director Chaos Engineering jobs in Ridgewood, NJ are:

What cities near Ridgewood, NJ are hiring for Director Chaos Engineering jobs?

Cities near Ridgewood, NJ with the most Director Chaos Engineering job openings:

Senior Production Engineer

CoreWeave

New York, NY • On-site

$139K - $204K/yr

Full-time

Re-posted 23 days ago


CoreWeave rating

9.8

Company rating: 9.8 out of 10

Based on 7 frontline employees who took The Breakroom Quiz

1st of 224 rated it services


Job description

About the Role

Production Engineering ensures CoreWeave's cloud delivers world-class reliability, performance, and operational excellence. We are hiring a Senior Production Engineer to take direct, hands-on ownership of critical tooling that drives reliability and delivery success.

In this role, you will work broadly across the cloud stack designing, implementing, deploying, and operating systems that improve delivery velocity, service availability, and operational safety. You'll be responsible for leading end-to-end technical projects, maintaining long-lived systems the team owns, and strengthening our operational foundations through durable engineering investments.

This is a role for someone who enjoys buildingdebugging, and operating production systems. You will collaborate closely with service owners, but your primary impact comes from the reliability, quality, and maturity of the systems you deliver and maintain over time.

What You'll Do
  • Take hands-on ownership of critical systems and frameworks, driving their architecture, implementation, and long-term evolution.
  • Lead end-to-end delivery of engineering projects that improve availability, scalability, operational automation, and failure recovery.
  • Build and maintain observability, alerting, automated remediation, and resilience testing for the systems you support.
  • Participate in incident response as a subject-matter expert; drive deep root-cause investigations and implement lasting fixes.
  • Improve runbooks, sources of truth, deployment workflows, and operational tooling to harden production readiness.
  • Eliminate single points of failure and reduce operational toil through automation, refactors, and system redesigns.
  • Ship production code regularly in Python, Go, or similar languages, and participate in on-call rotations.
  • Maintain and mature long-term projects and frameworks owned by the team, ensuring they remain reliable, well-instrumented, and easy to operate.
  • Collaborate with platform teams to ensure new features and services integrate cleanly with our reliability best-practices and tooling.
What You've Worked On (Minimum Qualifications)
  • 7+ years of engineering experience building and operating distributed systems or cloud platforms.
  • Demonstrated ability to debug complex production issues end-to-end, across services, infrastructure layers, and automation.
  • Strong programming or scripting ability (Python, Go, or similar), with experience shipping and operating production services and tools.
  • Deep knowledge of cloud-native technologies and distributed system patterns, particularly Kubernetes.
  • Experience with modern observability stacks: metrics, tracing, structured logs, SLOs/SLIs, and incident lifecycle practices.
  • A track record of successfully delivering hands-on reliability improvements through engineering execution.
Preferred Qualifications
  • Experience building internal tooling, frameworks, or automation that supports high-availability cloud operations.
  • Familiarity with DR/BCP, service tiering, capacity planning, or chaos engineering.
  • Background operating or building large-scale AI or GPU-accelerated infrastructure.
  • Experience maintaining multi-year ownership of foundational production systems.

Why CoreWeave?

At CoreWeave, we work hard, have fun, and move fast!  We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values: 

  • Be Curious at Your Core
  • Act Like an Owner
  • Empower Employees
  • Deliver Best-in-Class Client Experiences
  • Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the organization's growth opportunities are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us! 

The base salary range for this role is $139,000 to $204,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility). 


What CoreWeave employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom