1

Director Observability Jobs in Madison, WI (NOW HIRING)

This is a senior technical leadership role for someone who is exceptionally self-directed, highly ... observability, security, and cost optimisation * Shape the design and optimisation of data systems ...

New

... and observability. For more information, visit www.enterprisedb.com As EDB Principal Customer ... Direct experience working with globally distributed and remote customer and internal teams

Director Observability information

What is the difference between Director Observability vs Site Reliability Engineer?

AspectDirector ObservabilitySite Reliability Engineer
Primary FocusOversees observability strategies, tools, and teams to ensure system visibility and performanceBuilds and maintains reliable systems, automates deployment, and manages incident response
CredentialsTypically requires advanced knowledge of monitoring, cloud platforms, and leadership experienceOften has software engineering background, with skills in scripting, automation, and systems engineering
Work EnvironmentLeads teams in tech companies, focusing on monitoring and analytics toolsWorks closely with development and operations teams to ensure system reliability

While both roles focus on system performance and reliability, the Director Observability primarily manages observability strategies and teams, whereas the Site Reliability Engineer is hands-on, building and maintaining reliable systems. The roles complement each other in ensuring optimal system performance and uptime.

What are the key skills and qualifications needed to thrive as a director of observability, and why are they important?

To thrive as a Director of Observability, you need deep expertise in monitoring, logging, and distributed systems, typically backed by a degree in computer science or a related field and extensive experience in IT or DevOps leadership roles. Proficiency with observability tools such as Prometheus, Grafana, Datadog, Splunk, and APM solutions, along with knowledge of cloud platforms and relevant certifications, is essential. Strong leadership, strategic thinking, and communication skills help drive cross-functional initiatives and foster a culture of reliability. These skills and qualities are crucial for ensuring system health, rapid incident response, and alignment between technical teams and organizational objectives.

How does a director of observability typically collaborate with engineering and operations teams to drive organizational goals?

A Director of Observability works closely with engineering and operations teams to ensure that systems are monitored effectively and issues are identified and resolved quickly. This collaboration often involves developing unified monitoring strategies, aligning observability tools and processes, and facilitating incident response post-mortems. The Director also leads cross-functional meetings to establish best practices, set key performance indicators (KPIs), and ensure observability is integrated into the software development lifecycle. By acting as a bridge between technical teams, they help foster a culture of transparency, reliability, and continuous improvement.

What does a director of observability do?

A Director of Observability leads the strategy and implementation of monitoring, logging, and tracing systems to ensure the health and performance of technical infrastructure. They work with engineering and operations teams to develop best practices, select appropriate tools, and set standards for observability across the organization. Their goal is to provide visibility into system behavior, quickly identify and resolve incidents, and support continuous improvement in system reliability and performance.
What are popular job titles related to Director Observability jobs in Madison, WI? For Director Observability jobs in Madison, WI, the most frequently searched job titles are:
What job categories do people searching Director Observability jobs in Madison, WI look for? The top searched job categories for Director Observability jobs in Madison, WI are:
Infographic showing various Director Observability job openings in Madison, WI as of August 2026, with employment types broken down into 96% Full Time, and 4% Contract. Highlights an 82% In-person, and 18% Remote job distribution.

AI Platform Engineering Director (Primarily Office)

American Family Mutual Insurance Company Si

Madison, WI • On-site

$172K - $294K/yr

Full-time

Medical, Dental, Vision, Retirement, PTO

Posted 21 days ago


American Family Insurance rating

7.5

Company rating: 7.5 out of 10

Based on 135 frontline employees who took The Breakroom Quiz

217th of 304 rated insurance


Job description

This position provides strategic and technical leadership for the enterprise AI platform engineering function, setting the technology strategy and reference architecture for how AI is built, deployed, and scaled enterprise-wide. The role is accountable for building and operating shared AI engineering capabilities that enable teams to develop, deploy, monitor, evaluate, and scale AI solutions safely and efficiently. As Director, you will lead teams responsible for AI platform architecture, reusable engineering frameworks, MLOps/LLMOps, AI operations, agentic workflow patterns, LLM pipeline frameworks, observability, evaluation controls, production support patterns, systems-of-record integration, cost optimization, and technical guardrails that support responsible and scalable AI adoption. This role works across Technology and business domains, including Information Security, Governance, Enterprise Architecture, Application Development, Digital Services, Infrastructure, Data Engineering, Product, Operations, Risk, Legal, Compliance, and strategic technology partners. You will also manageor coordinates key AI platform partners, including GCP, AWS, DataDog, ServiceNow, Salesforce and more where applicable.

Position Compensation Range:

$172,000.00 - $294,000.00

Pay Rate Type:

Salary

Compensation may vary based on the job level and your geographic work location. Relocation support is offered for eligible candidates.

Primary Accountabilities

  • Lead AI platform architecture and strategy
    • You will define the architecture, standards, and roadmap for shared enterprise AI platform capabilities. This is collaborative with Data & AI Architects.
    • You will also ensure the platform supports scalable, secure, reliable, and governed AI delivery across multiple business domains.
  • Balance AI operations, MLOps/LLMOps, and agentic frameworks
    • Lead engineering practices for model and AI-system deployment, monitoring, testing, evaluation, versioning, reliability, and lifecycle management.
    • Balance operational reliability with reusable agentic workflow patterns and LLM pipeline frameworks.
    • Ensure AI systems can be supported and improved after production deployment.
  • Contribute to the enterprise AI technology maturity view
    • You will contribute platform, engineering, operations, observability, support, resilience, cost, systems integration, and production-readiness inputs to the enterprise-level shared AI technology maturity view.
    • You will use this maturity view to identify capability gaps, guide investment recommendations, and communicate platform and engineering readiness.
  • Build reusable engineering frameworks and capabilities
    • Build and maintain reusable frameworks, components, and patterns that accelerate AI delivery and reduce duplicated engineering effort.
    • Ensure durable reusable capabilities are documented, discoverable, supportable, and governed so they can be leveraged across multiple domains.
  • Own platform production support and operational run patterns
    • You will directly own production support for the AI platform and shared AI engineering capabilities, especially L1, L2, and the engineering side of L3 support
    • Establish production run patterns for AI-enabled workflows, including support models, incident paths, escalation patterns, fallback mechanisms, and human handoff design.
    • Partner with customer service, employee support, application support, service management, or other operational channels where AI experiences require context-rich support transitions.
  • Establish observability, evaluation, and monitoring capabilities
    • Establish platform capabilities for telemetry, model and system monitoring, traceability, drift or quality signals, evaluation controls, regression checks, and operational visibility.
    • Provide visibility into AI system behavior, performance, usage, and risk signals.
  • Embed responsible AI technical guardrails
    • Embed responsible AI technical controls into platform capabilities in alignment with shared enterprise governance expectations.
    • Ensure platform capabilities support auditability, policy adherence, access controls, least-privilege operation, model/agent registration, and responsible AI practices.
  • Lead AI cost, capacity, and resilience practices
    • Provide engineering mechanisms for token strategy, usage monitoring, multi-model routing, fallback, caching, capacity planning, inference or compute optimization, and total-cost-of-ownership discipline.
    • Help the enterprise balance speed, performance, reliability, resilience, and cost.
  • Support build/buy/partner technical decisions
    • Lead build/buy/partner decisions for AI engineering capabilities, including when to use partner-native agents, managed AI platforms, or internally built orchestration based on data location, control needs, governance, speed, cost, portability, and strategic differentiation.
    • Manage and coordinate key AI Platform partners, including GCP, AWS, DataDog, ServiceNow, and others where applicable.
  • Lead people and develop engineering talent
    • Lead teams responsible for MLOps, LLMOps, AI operations, platform engineering, GIS or other assigned platform capabilities, and related AI engineering functions.
    • Build engineering discipline, technical depth, delivery accountability, and collaborative execution across teams, while recognizing the AI space is evolving quickly and required skills will continue to evolve.

Specialized Knowledge & Skills Requirements

  • Demonstrated experience leading engineering teams that build production platforms, internal developer platforms, MLOps/LLMOps capabilities, AI operations, or scalable AI/ML systems.
  • Experience with AI system architecture, model deployment, agent deployment, monitoring, observability, evaluation, production support, and lifecycle management.
  • Demonstrated ability to balance operational reliability with emerging agentic workflow and LLM pipeline frameworks.
  • Experience creating reusable frameworks, standards, and platform capabilities that improve delivery across multiple teams.
  • Familiarity with GenAI, agentic workflows, LLM pipelines, model orchestration, retrieval-augmented generation patterns, model gateways, systems-of-record integration, and emerging AI platform patterns.
  • Demonstrated experience with production support models, including L1/L2 support expectations and engineering-side L3 support.
  • Experience with cost, capacity, resilience, usage monitoring, routing, fallback, caching, or FinOps practices for cloud or AI workloads.
  • Experience working across Information Security, Enterprise Architecture, Infrastructure, Cloud, Application Development, Digital Services, Data Engineering, Governance, Legal, Risk, Compliance, and business domains.
  • Experience managing or coordinating partners such as GCP, AWS, DataDog, or related technology vendors.
  • Demonstrated people leadership, technical coaching, prioritization, and talent development skills.

Key Interfaces and Partners

  • Applied AI for solution delivery needs, reusable patterns, production enablement, applied feedback loops, L3 enhancement partnership, and solution handoff.
  • BI Engineering & Enablement for metric, semantic, metadata, lineage, and knowledge-layer dependencies that support AI workflows.
  • Data Engineering for data pipelines, data products, environment dependencies, data availability, and data readiness.
  • Information Security, Enterprise Architecture, Infrastructure, Cloud, SRE, Application Development, Digital Services, Privacy, Legal, Compliance, Model Risk, AI Governance, Data Governance, and Procurement/TPRO.
  • Product and business-domain teams using AI platform capabilities.
  • Finance or FinOps partners for AI cost visibility and optimization.
  • Strategic technology partners and platform teams supporting Salesforce, ServiceNow, Guidewire, Workday, GCP, AWS, Microsoft, Google, DataDog, and other AI ecosystem components.
Additional Information
  • To ensure a strong start, all employees participate in our New Employee Orientation during their first week. This experience is held in person at our Madison, WI Headquarters or one of our AmFam core locations to help you connect with our mission, meet key team members and build relationships that support your growth. At times, sessions may be delivered virtually based on scheduling and availability.
  • Offer to selected candidate will be made contingent on the results of applicable background checks
  • Offer to selected candidate is contingent on signing a non-disclosure agreement for proprietary information, trade secrets, and inventions
  • Sponsorship will not be considered for this position unless specified in the posting

In this primarily office-based role, you will be expected to spend at least 80% of your time (4+ days per week) working from the office. Candidates should reside within approximately 35-50 miles of one of the following office locations: Madison, WI 53783; or Boston, MA 02110.

#LI-Onsite

We provide benefits that support your physical, emotional, and financial wellbeing. You will have access to comprehensive medical, dental, vision and wellbeing benefits that enable you to take care of your health. We also offer a competitive 401(k) contribution, a pension plan, an annual incentive, 9 paid holidays and a paid time off program (23 days accrued annually for full-time employees). In addition, our student loan repayment program and paid-family leave are available to support our employees and their families. Interns and contingent workers are not eligible for American Family Insurance Group benefits.

We are an equal opportunity employer. It is our policy to comply with all applicable federal, state and local laws pertaining to non-discrimination, non-harassment and equal opportunity. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law.

American Family Insurance is committed to the full inclusion of all qualified individuals. If a reasonable accommodation is needed to participate in the job application or interview process, to perform essential job functions, and/or to receive other benefits and privileges of employment, please email AskHR@AmFam.com to request a reasonable accommodation.

#LI-AW1

What American Family Insurance employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom