1

Director Observability Jobs in Michigan (NOW HIRING)

Sr. AI Engineer

Southfield, MI · On-site

$95K - $131K/yr

Strong software engineering fundamentals - testing, CI/CD, observability, error handling - applied ... direct line to the Director of APEX. The platform is early - that is the appeal. You are not ...

Sr. AI Engineer

Southfield, MI

$95K - $131K/yr

... the observability layer, and the deployment infrastructure around the models. Building that ... direct line to the Director of APEX. The platform is early - that is the appeal. You are not ...

next page

Showing results 1-20

Director Observability information

What is the difference between Director Observability vs Site Reliability Engineer?

AspectDirector ObservabilitySite Reliability Engineer
Primary FocusOversees observability strategies, tools, and teams to ensure system visibility and performanceBuilds and maintains reliable systems, automates deployment, and manages incident response
CredentialsTypically requires advanced knowledge of monitoring, cloud platforms, and leadership experienceOften has software engineering background, with skills in scripting, automation, and systems engineering
Work EnvironmentLeads teams in tech companies, focusing on monitoring and analytics toolsWorks closely with development and operations teams to ensure system reliability

While both roles focus on system performance and reliability, the Director Observability primarily manages observability strategies and teams, whereas the Site Reliability Engineer is hands-on, building and maintaining reliable systems. The roles complement each other in ensuring optimal system performance and uptime.

What are the key skills and qualifications needed to thrive as a director of observability, and why are they important?

To thrive as a Director of Observability, you need deep expertise in monitoring, logging, and distributed systems, typically backed by a degree in computer science or a related field and extensive experience in IT or DevOps leadership roles. Proficiency with observability tools such as Prometheus, Grafana, Datadog, Splunk, and APM solutions, along with knowledge of cloud platforms and relevant certifications, is essential. Strong leadership, strategic thinking, and communication skills help drive cross-functional initiatives and foster a culture of reliability. These skills and qualities are crucial for ensuring system health, rapid incident response, and alignment between technical teams and organizational objectives.

How does a director of observability typically collaborate with engineering and operations teams to drive organizational goals?

A Director of Observability works closely with engineering and operations teams to ensure that systems are monitored effectively and issues are identified and resolved quickly. This collaboration often involves developing unified monitoring strategies, aligning observability tools and processes, and facilitating incident response post-mortems. The Director also leads cross-functional meetings to establish best practices, set key performance indicators (KPIs), and ensure observability is integrated into the software development lifecycle. By acting as a bridge between technical teams, they help foster a culture of transparency, reliability, and continuous improvement.

What does a director of observability do?

A Director of Observability leads the strategy and implementation of monitoring, logging, and tracing systems to ensure the health and performance of technical infrastructure. They work with engineering and operations teams to develop best practices, select appropriate tools, and set standards for observability across the organization. Their goal is to provide visibility into system behavior, quickly identify and resolve incidents, and support continuous improvement in system reliability and performance.

What are the most commonly searched types of Observability jobs in Michigan?

The most popular types of Observability jobs in Michigan are:

What are popular job titles related to Director Observability jobs in Michigan?

For Director Observability jobs in Michigan, the most frequently searched job titles are:

What cities in Michigan are hiring for Director Observability jobs?

Cities in Michigan with the most Director Observability job openings:

$180 - $240/hr

Other

Posted 3 days ago

New


Papa John's rating

4.8

Company rating: 4.8 out of 10

Based on 754 frontline employees who took The Breakroom Quiz

19th of 24 rated food delivery companies


Job description

Job Summary

The Director of Technology Enablement is responsible for the daily operational stability, service performance, and execution of core technology operations across cloud, infrastructure, networks, enterprise platforms, and service delivery. This role ensures production environments remain highly available, fully supported, secure, and compliant, and that operational processes run with discipline, repeatability, and measurable outcomes. This role will also serve as a SME liaison across all the Technology and Infrastructure group, helping to maximize collaboration beginning with all the foundational compute layout (architecture and networking), to enterprise platform design and build phases, all the way through site reliability and monitoring and observability after production implementations.

Duties and Responsibilities (other duties as assigned) Operational Command & Production Stability
  • Own 24/7 operational responsibility for all production systems, platforms, and services.
  • Establish and manage daily operational health checks, system dashboards, and service reviews.
  • Lead Major Incident Management (MIM) with defined escalation paths, communicator roles, and incident commander playbooks.
  • Maintain runbooks, failover plans, and operational procedures for all critical services.
Service Delivery & Support Operations
  • Manage Tier 1–3 operational support teams across infrastructure, cloud, network, SRE/DevOps, and application operations.
  • Execute on SLA/OLA frameworks and enforce response/resolution timelines.
  • Oversee service desk operations including ITSM workflows (incidents, requests, changes, problems, knowledge).
  • Ensure operational readiness for new applications, releases, or infrastructure deployments.
Change, Release, & Deployment Operations
  • Help lead the Change Advisory Board (CAB) weekly sessions and enforce change management guardrails.
  • Ensure controlled release pipelines, deployment windows, rollback procedures, and post-change validation steps.
  • Maintain a controlled maintenance calendar coordinating across all service owners.
Incident, Problem, and RCA Execution
  • Enforce clear MTTA/MTTR targets for high-severity events.
  • Oversee Root Cause Analysis (RCA) processes, ensuring corrective actions and preventative measures are completed.
  • Track incident trends, repeat failures, and configuration drift.
  • Oversee problem-management routines and drive long-term stability improvements.
Monitoring, Observability & Operational Analytics
  • Manage real-time operations dashboards: availability, latency, saturation, capacity, and error rates.
  • Ensure proactive alerting, threshold tuning, false-positive reduction, and early-warning detection.
  • Run weekly operations review meetings with service owners using KPI-driven analysis.
  • Drive capacity monitoring across cloud and data center environments (compute, storage, network).
Infrastructure & Cloud Operations Execution
  • Ensure patching schedules, backup success rates, restore testing, and disaster recovery drills are consistently executed.
  • Govern configuration management, standard builds, and environment consistency across dev/test/prod.
  • Manage provisioning, resource tagging, lease expiration, cost monitoring, and capacity allocation.
  • Maintain operational hygiene: certificate renewals, OS lifecycle management, etc.
  • Add operational POV to all architectural, networking, storage, and information lifecycle management elements of our PJI technology ecosystem.
Operational Governance, Compliance, & Risk Management
  • Own operational controls related to security, compliance, and audit readiness.
  • Ensure completion of vulnerability remediation cycles, compliance evidence, and operational control testing.
  • Maintain DR runbooks, RTO/RPO compliance, and oversee annual or semiannual DR failover exercises.
  • Maintain asset inventories, CMDB accuracy, and configuration baselines.
  • Lead technology FinOps practice to streamline and maximize our annual budgetary commitments and vendor partnerships.
Workforce Management & Operational Performance
  • Manage shift scheduling, on-call rotations, follow-the-sun support models, and operational staffing levels.
  • Ensure operational training, certification, and cross-skilling of team members.
  • Run operational playbook updates, SOP reviews, and continuous improvement initiatives.
Education, Experience & Certifications
  • Bachelor’s degree in Information Technology, Computer Science, or related field
  • Relevant certifications such as ITIL, CompTIA, or cloud platform certifications (AWS, Azure, or Google Cloud) preferred
  • Proven experience in IT operations, service management, or infrastructure support
  • Strong understanding of SLOs/SLAs and incident management best practices
  • Experience with change management, patch management, and vulnerability remediation
Functional Skills
  • Monitoring and ensuring service availability according to SLO/SLA requirements
  • Incident detection, triage, and resolution with focus on reducing MTTD/MTTR
  • Implementing and tracking changes, analyzing change success rates, and minimizing change-related incidents
  • Root cause analysis to prevent incident recurrence
  • Managing patch compliance and closing vulnerabilities within defined timelines
  • Coordinating and validating backups and disaster recovery tests
  • Optimizing operational costs and managing cloud/infrastructure spending
  • Fulfilling service requests efficiently and maintaining accurate operational documentation (runbooks, SOPs, diagrams, CMDB)
  • Planning strategy for compute capacity growth and expansion needs
Our Values
  • EVERYONE BELONGS – We believe connectedness and belonging are the essential ingredients to our success.
  • DO THE RIGHT THING –We are relentlessly focused on quality and integrity and make the right choices, even when it's difficult.
  • PEOPLE FIRST – To craft positive experiences for our customers, we take care of each other first.
  • INNOVATE TO WIN – We champion and challenge for a better way in all we do.
  • HAVE FUN – We find joy, create meaningful impact and celebrate the journey together
Our Core Competencies
  • CUSTOMER CENTRIC - We leverage data and insights to craft a customer experience that builds relationships, cultivates trust, and delivers excellence
  • RESULTS DRIVEN – We focus on measurable outcomes by remaining optimistic, tenacious, and persistent even in the face of challenges.
  • CONTINUOUS IMPROVEMENT –We champion for better through strategic risk taking, experimentation and challenging the status quo.
  • BIAS FOR ACTION – We courageously lead, drive towards decisions, and maintain agility to meet the demands of our dynamic industry.
  • WINNING TOGETHER – We work together to unlock our full potential by actively collaborating and contributing in a cross-functional capacity

Papa Johns is an equal opportunity employer.

Papa Johns is a federal contractor that participates in the E-Verify program to confirm employment eligibility for each new team member. We also comply with all Right to Work requirements. Official E-Verify and Right to Work notices are available for applicants to review in both English and Spanish.

#J-18808-Ljbffr

What Papa John's employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom