1

Director Observability Jobs (NOW HIRING)

Senior Director - Observability | SRE

Coppell, TX · On-site

$52.50 - $70/hr

About the Role The Senior Director - Observability and SRE, is a strategic leader accountable for ensuring the reliability, availability, and performance of the enterprise technology ecosystem. This ...

Senior Director - Observability | SRE

Coppell, TX · On-site

$53 - $70.50/hr

About the role The Senior Director - Observability and SRE, is a strategic leader accountable for ensuring the reliability, availability, and performance of the enterprise technology ecosystem. This ...

We are seeking a Director of Observability to stand up a brand-new observability and reliability practice from the ground up at Kestra Holdings. This is a working leadership role - the Director will ...

$49.76 - $82.94/hr

Wir bieten dir ein modernes Arbeitsumfeld mit flachen Hierarchien, agilen Arbeitsmethoden und Raum für neue Ideen - Du hast vielfältige Möglichkeiten zur Gestaltung unserer Unternehmenskultur - Du ...

next page

Showing results 1-20

Director Observability information

What does a director of observability do?

A Director of Observability leads the strategy and implementation of monitoring, logging, and tracing systems to ensure the health and performance of technical infrastructure. They work with engineering and operations teams to develop best practices, select appropriate tools, and set standards for observability across the organization. Their goal is to provide visibility into system behavior, quickly identify and resolve incidents, and support continuous improvement in system reliability and performance.

How does a director of observability typically collaborate with engineering and operations teams to drive organizational goals?

A Director of Observability works closely with engineering and operations teams to ensure that systems are monitored effectively and issues are identified and resolved quickly. This collaboration often involves developing unified monitoring strategies, aligning observability tools and processes, and facilitating incident response post-mortems. The Director also leads cross-functional meetings to establish best practices, set key performance indicators (KPIs), and ensure observability is integrated into the software development lifecycle. By acting as a bridge between technical teams, they help foster a culture of transparency, reliability, and continuous improvement.

What are the key skills and qualifications needed to thrive as a director of observability, and why are they important?

To thrive as a Director of Observability, you need deep expertise in monitoring, logging, and distributed systems, typically backed by a degree in computer science or a related field and extensive experience in IT or DevOps leadership roles. Proficiency with observability tools such as Prometheus, Grafana, Datadog, Splunk, and APM solutions, along with knowledge of cloud platforms and relevant certifications, is essential. Strong leadership, strategic thinking, and communication skills help drive cross-functional initiatives and foster a culture of reliability. These skills and qualities are crucial for ensuring system health, rapid incident response, and alignment between technical teams and organizational objectives.

What is the difference between Director Observability vs Site Reliability Engineer?

AspectDirector ObservabilitySite Reliability Engineer
Primary FocusOversees observability strategies, tools, and teams to ensure system visibility and performanceBuilds and maintains reliable systems, automates deployment, and manages incident response
CredentialsTypically requires advanced knowledge of monitoring, cloud platforms, and leadership experienceOften has software engineering background, with skills in scripting, automation, and systems engineering
Work EnvironmentLeads teams in tech companies, focusing on monitoring and analytics toolsWorks closely with development and operations teams to ensure system reliability

While both roles focus on system performance and reliability, the Director Observability primarily manages observability strategies and teams, whereas the Site Reliability Engineer is hands-on, building and maintaining reliable systems. The roles complement each other in ensuring optimal system performance and uptime.

More about Director Observability jobs

What cities are hiring for Director Observability jobs?

Cities with the most Director Observability job openings:

What are the most commonly searched types of Observability jobs?

The most popular types of Observability jobs are:

What states have the most Director Observability jobs?

States with the most job openings for Director Observability jobs include:

Infographic showing various Director Observability job openings in the United States as of August 2026, with employment types broken down into 100% Full Time. Highlights an 75% In-person, and 25% Remote job distribution.

Senior Director - Observability | SRE

Gap, Inc.

Coppell, TX • On-site

$52.50 - $70/hr

Full-time

Posted 18 days ago


Gap rating

6.8

Company rating: 6.8 out of 10

Based on 279 frontline employees who took The Breakroom Quiz

29th of 105 rated fashion retailers


Job description

About the Role
The Senior Director - Observability and SRE, is a strategic leader accountable for ensuring the reliability, availability, and performance of the enterprise technology ecosystem. This role oversees Observability, Site Reliability Engineering (SRE), Visibility and Live Sight Insights. This leader drives operational excellence through a proactive strategy that combines process discipline, automation, observability, and real-time insights. They will partner closely with engineering, infrastructure, cybersecurity, and product teams to build and sustain systems that power Gap Inc.'s digital and in-store experiences. As a thought leader, the Sr. Director will shape the long-term vision for operational reliability, defining modern capabilities, optimizing service performance, and establishing an innovation-driven reliability culture.What You'll Do
Strategic Leadership & Vision
  • Define and execute the enterprise Observability and SRE strategy, ensuring alignment with business objectives and technology roadmaps.
  • Lead transformation of Unified Observability for end-to-end visibility of systems to actively and proactively reduce mean time to resolve through critical path anomaly detection.
  • Partner with senior technology and business leaders to embed reliability and performance metrics into product development and operational planning.

Operational Excellence & Reliability Engineering
  • Lead Site Reliability Engineering (SRE) practices across platforms and services driving automation, self-healing capabilities, and proactive monitoring to achieve measurable service resiliency improvements.
  • Establish standards for availability, latency, scalability, and operational efficiency through engineering-driven reliability principles.
  • Champion reliability by design ensuring observability, capacity planning, and chaos testing are core to delivery processes.

Mission Control & Live Sight Insights
  • Oversee the Mission Control organization responsible for real time system monitoring, across ecommerce, fulfillment centers, and stores.
  • Drive adoption of Live Sight Insights to create predictive and actionable intelligence on service health and performance trends.
  • Enable enterprise visibility of key metrics through intuitive dashboards and business-impact-based alerting models.
  • Lead a platform governance mindset focusing on reliability, scalability, and ease of use.

People Leadership & Culture
  • Build, inspire, and develop a high-performing global Observability and SRE team that embodies accountability, collaboration, and innovation.
  • Foster a culture of data driven decision making, continuous learning, and operational excellence.
  • Serve as a mentor and coach to emerging leaders raising the organizational bar for reliability engineering and service leadership.

Cross-Functional Partnership
  • Work closely with Software Engineering, Infrastructure, Cybersecurity, and Business Technology teams to ensure reliability objectives are integrated end-to-end.
  • Partner with Enterprise Architecture and Program Management to align technology investments with reliability outcomes.
  • Act as a trusted advisor to executive leadership on reliability strategy, risk posture, and enterprise service health and performance.
Who You Are
  • Proven strategic leader with success driving operational transformation at scale in global, complex environments for more than 10 years.
  • Deep expertise in ITIL frameworks, SRE principles, and architecture, and modern observability and SRE practices.
  • Strong technical understanding across infrastructure, cloud operations, automation, and service management ecosystems.
  • Exceptional ability to influence at all levels translating technical reliability concepts into business impact and strategic value.
  • Passionate about developing people and creating a culture of ownership, reliability, and continuous improvement.
  • Demonstrated track record of leading large, diverse teams and delivering measurable improvements in service reliability, performance, and user satisfaction.
  • A high performing leader operating with strategic agility, executive presence, and the ability to build organizational alignment through clarity, accountability, and purpose.

What Gap employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom