1

Executive Observability Engineer Jobs (NOW HIRING)

Data & Observability Architect

Dallas, TX · On-site

$200K - $325K/yr

Help build persona-oriented views for Finance, Operation, Executives, Developers, Platform etc. * Build and guide transparency around cost, observability and resiliency of the observability platform.

Background in SRE or reliability engineering, with strong understanding of SLOs, incident ... S. law, regulation, executive order, or government contract. For Los Angeles County (unincorporated ...

Background in SRE or reliability engineering, with strong understanding of SLOs, incident ... S. law, regulation, executive order, or government contract. For Los Angeles County (unincorporated ...

The Opportunity Dash0 is looking for an Enterprise Account Executive to spearhead our entry into ... Experience in observability, DevOps, cloud infrastructure, or data platforms. * Existing ...

Senior Grafana Engineer

New York, NY · On-site

$82K - $193K/yr

Executive Summary The Senior Grafana Engineer will be responsible for enterprise observability platform engineering, Grafana Cloud migration, OpenTelemetry adoption, automation, integrations ...

next page

Showing results 1-20

Executive Observability Engineer information

What is the difference between Executive Observability Engineer vs Site Reliability Engineer?

AspectExecutive Observability EngineerSite Reliability Engineer
CredentialsTypically requires expertise in observability tools, monitoring, and cloud platformsRequires skills in systems engineering, coding, and infrastructure management
Work EnvironmentFocuses on designing observability solutions, analyzing system health, and strategic monitoringManages system reliability, automates deployment, and maintains infrastructure
Industry UsageUsed in tech companies emphasizing system visibility and performance analysisCommon in cloud services, SaaS, and large-scale web services

The Executive Observability Engineer primarily focuses on implementing and optimizing observability tools to ensure system health, while the Site Reliability Engineer concentrates on maintaining system reliability and automating infrastructure. Both roles require technical expertise but differ in their strategic versus operational focus.

What are some common challenges Executive Observability Engineers face when implementing organization-wide monitoring solutions?

Executive Observability Engineers often encounter challenges such as integrating diverse monitoring tools across legacy and modern systems, ensuring data consistency, and balancing comprehensive visibility with system performance. Coordinating with multiple teams to align observability goals and fostering a culture of proactive monitoring can also require strong communication and leadership skills. Successfully managing these complexities leads to more resilient infrastructure and improved incident response times.

What are Executive Observability Engineers?

Executive Observability Engineers are specialized IT professionals who design and manage systems that monitor, analyze, and optimize the health and performance of an organization's technology infrastructure. They focus on providing high-level visibility into applications, networks, and services, enabling executives to make data-driven decisions. These engineers implement observability tools, create dashboards, and generate reports that translate complex technical metrics into actionable business insights. Their work is crucial for ensuring system reliability, quick incident response, and continuous improvement across the organization.

What are the key skills and qualifications needed to thrive as an Executive Observability Engineer, and why are they important?

To thrive as an Executive Observability Engineer, you need deep expertise in performance monitoring, distributed systems, and troubleshooting, often supported by a degree in computer science or a related field. Familiarity with observability tools such as Datadog, New Relic, Prometheus, and advanced logging systems, as well as certifications like AWS Certified DevOps Engineer, is highly beneficial. Strong analytical thinking, communication, and leadership skills set top candidates apart in this role. These skills and qualities are crucial to proactively detect issues, optimize system performance, and drive strategic decision-making across complex technical environments.
More about Executive Observability Engineer jobs
What cities are hiring for Executive Observability Engineer jobs? Cities with the most Executive Observability Engineer job openings:
What are the most commonly searched types of Observability Engineer jobs? The most popular types of Observability Engineer jobs are:
What states have the most Executive Observability Engineer jobs? States with the most job openings for Executive Observability Engineer jobs include:
What job categories do people searching Executive Observability Engineer jobs look for? The top searched job categories for Executive Observability Engineer jobs are:
Infographic showing various Executive Observability Engineer job openings in the United States as of July 2026, with employment types broken down into 96% Full Time, 1% Part Time, and 3% Contract. Highlights an 87% Physical, 5% Hybrid, and 8% Remote job distribution.
Senior Site Reliability Engineer - Unified Observability

Senior Site Reliability Engineer - Unified Observability

NCR

Atlanta, GA • On-site

$54.75 - $72.75/hr

Full-time

Posted 2 days ago


NCR Corporation rating

6.6

Company rating: 6.6 out of 10

Based on 10 frontline employees who took The Breakroom Quiz

183rd of 215 rated software companies


Job description

About NCR VOYIX

NCR Voyix Corporation (NYSE: VYX) is a global platform-powered leader in unified commerce for shopping and dining. Combining a flexible, intelligent platform with end-to-end payments capabilities and services developed through its deep industry experience, NCR Voyix empowers retailers and restaurants to accelerate new possibilities for their operations, experiences and business outcomes. NCR Voyix is headquartered in Atlanta, Georgia, and serves customers in more than 35 countries worldwide.
Position Overview

We are seeking a highly experienced Senior Site Reliability Engineer (Unified Observability) to lead the design, implementation, and operational maturity of the F1 Next Generation Customer Unified Observability initiative. This strategic role will be responsible for building and evolving a unified enterprise observability platform that delivers end-to-end visibility across NCR Voyix Restaurants, Retail, and Payments environments.

The ideal candidate will bring 10+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, or related disciplines, with a proven track record of driving enterprise-scale observability, reliability, and operational excellence. This individual must be comfortable operating across organizational boundaries and partnering closely with Product Engineering, Infrastructure, Security, Operations, Architecture, and Executive Leadership teams to establish a comprehensive observability strategy and improve platform resilience.

This role will serve as a key technical leader responsible for defining standards, influencing architecture decisions, and enabling proactive operations through unified monitoring, telemetry, automation, and AI-driven insights.

Key Responsibilities
  • Lead the architecture, design, implementation, and continuous improvement of enterprise observability solutions across Azure, Google Cloud Platform (GCP), Kubernetes, and hybrid environments.
  • Establish and drive enterprise observability standards for monitoring, logging, distributed tracing, telemetry, and operational analytics.
  • Develop and maintain executive, operational, and engineering dashboards that provide real-time visibility into infrastructure, applications, platform health, customer experience, and business transactions.
  • Define, evangelize, and implement reliability frameworks including SLIs, SLOs, error budgets, operational KPIs, and service health metrics.
  • Partner cross-functionally with Engineering, Infrastructure, Security, Product, and Operations teams to identify reliability risks and drive operational excellence initiatives.
  • Lead efforts to improve incident prevention, detection, response, and recovery through intelligent alerting, automation, event correlation, and observability best practices.
  • Integrate observability capabilities with ServiceNow, CI/CD pipelines, automation frameworks, and enterprise operational workflows.
  • Influence technical strategy and roadmap decisions related to reliability engineering, platform observability, and operational readiness.
  • Support and drive enterprise initiatives involving AI-driven observability, predictive analytics, anomaly detection, and event intelligence.
  • Mentor engineers and serve as a subject matter expert for observability, reliability engineering, and cloud-native operations.
  • Establish governance, adoption, and best practices across multiple product and engineering teams to ensure consistent observability standards enterprise-wide.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
  • 10+ years of experience in Site Reliability Engineering, Cloud Engineering, Platform Engineering, DevOps, or related technical disciplines.
  • Demonstrated success designing and operating observability platforms in large-scale enterprise environments.
  • Deep expertise with Kubernetes platforms, including AKS and GKE.
  • Strong experience with Azure and Google Cloud Platform services and architectures.
  • Hands-on experience with enterprise observability tools such as Grafana, Datadog, Prometheus, OpenTelemetry, Dynatrace, New Relic, or similar platforms.
  • Advanced knowledge of monitoring, logging, telemetry collection, distributed tracing, and observability engineering principles.
  • Experience defining and operationalizing SLIs, SLOs, error budgets, reliability metrics, and service health frameworks.
  • Strong automation and Infrastructure as Code expertise using Terraform and related tools.
  • Proficiency developing automation solutions using Python, Go, PowerShell, or similar languages.
  • Experience integrating observability solutions into CI/CD pipelines and modern DevOps workflows.
  • Proven ability to influence technical direction and collaborate effectively with stakeholders across Engineering, Product, Infrastructure, Security, and Operations organizations.
  • Strong communication, leadership, and stakeholder management skills with the ability to translate technical concepts for both technical and business audiences.
Preferred Qualifications
  • Experience leading enterprise observability transformations or platform modernization initiatives.
  • Experience with AI Ops, event correlation, operational analytics, and predictive monitoring capabilities.
  • Knowledge of ServiceNow integrations and ITSM/ITOM processes.
  • Experience supporting highly available, customer-facing SaaS platforms at scale.
  • One or more cloud certifications (Azure, Google Cloud, Kubernetes, or related technologies).
  • Previous experience serving as a technical lead, mentor, or architect within a reliability engineering organization.

Offers of employment are conditional upon passage of screening criteria applicable to the job

EEO Statement

Integrated into our shared values is NCR Voyix's commitment to equal employment opportunity. All qualified applicants will receive consideration for employment without regard to sex, age, race, color, creed, religion, national origin, disability, sexual orientation, gender identity, veteran status, military service, genetic information, or any other characteristic or conduct protected by law. NCR Voyix is committed to being a globally inclusive company where all people are treated fairly, recognized for their individuality, promoted based on performance and encouraged to strive to reach their full potential. We believe in understanding and respecting differences among all people. Every individual at NCR Voyix has an ongoing responsibility to respect and support a globally diverse environment.

Statement to Third Party Agencies
To ALL recruitment agencies: NCR Voyix only accepts resumes from agencies on the preferred supplier list. Please do not forward resumes to our applicant tracking system, NCR Voyix employees, or any NCR Voyix facility. NCR Voyix is not responsible for any fees or charges associated with unsolicited resumes

"When applying for a job, please make sure to only open emails that you will receive during your application process that come from a @ncrvoyix.comemail domain."


What NCR Corporation employees say

Pay

Hours and flexibility

Workplace

Get the full story on Breakroom


NCR logo

About NCR

Sourced by ZipRecruiter

NCR Corporation is a leader in transforming, connecting and running technology platforms for self-directed banking, stores and restaurants. NCR is headquartered in Atlanta, Ga., with 38,000 employees globally. NCR is a trademark of NCR Corporation in the United States and other countries.

Industry

It services

Company size

10,000+ Employees

Headquarters location

Duluth, GA, US

Year founded

1884