1

Senior Observability Engineer Jobs in Detroit, MI

Senior Platform Engineer

Novi, MI ยท On-site

$117K/yr

As a Senior Platform Engineer, you will design, build, and operate cloud-native services that power ... Support AI-assisted observability initiatives, including intelligent alerting, anomaly detection ...

Sr. AI Engineer

Southfield, MI ยท On-site

$95K - $131K/yr

This is a deep, hands-on, senior individual-contributor role. You will own the implementation of ... Strong software engineering fundamentals - testing, CI/CD, observability, error handling - applied ...

Sr. AI Engineer

Southfield, MI ยท On-site

$95K - $131K/yr

This is a deep, hands-on, senior individual-contributor role. You will own the implementation of ... Strong software engineering fundamentals - testing, CI/CD, observability, error handling - applied ...

Senior Security Software Engineer

Warren, MI ยท On-site

$100 - $130/hr

The Role**As a senior engineer, you will lead the design, build, and operation of secure cloud ... Define and implement observability for services, including metrics, logs, traces, dashboards, and ...

Senior Security Software Engineer

Warren, MI ยท On-site

$115K - $151K/yr

The Role As a senior engineer, you will lead the design, build, and operation of secure cloud ... Define and implement observability for services, including metrics, logs, traces, dashboards, and ...

Core Senior Engineer

Dearborn, MI

$112K - $148K/yr

Software Engineer (3) - Core Senior Engineer #1058712 Position Description: * We are seeking a ... As a Software Engineer on our Platform Observability team, you will be at the heart of our ...

Senior Software Engineer

Auburn Hills, MI ยท On-site

$115K - $152K/yr

The Senior Software Engineer, Platform builds and operates the shared services, pipelines, and ... Establishes and improves the paved road: service templates, deployment pipelines, observability ...

Senior Software Engineer

Auburn Hills, MI ยท On-site

$115K - $152K/yr

The Senior Software Engineer, Platform builds and operates the shared services, pipelines, and ... Establishes and improves the paved road: service templates, deployment pipelines, observability ...

Senior Software Engineer

Birmingham, MI ยท On-site

$100 - $130/hr

Senior Software Engineer RPM is an international non-asset-based logistics and supply chain ... Spring Boot Actuator and AWS CloudWatch for health checks and observability * Experience with Auth0 ...

Senior Analytics Engineer

Detroit, MI ยท On-site

$103K - $142K/yr

Testing, Quality & Observability: Establish and enforce testing and documentation standards in dbt ... engineering, or ELT/data modeling, including demonstrated senior-level ownership of a ...

Senior Platform Engineer

Ann Arbor, MI ยท On-site

$102K - $140K/yr

We are seeking a Senior / Lead Platform Engineer to design, build, and evolve a highly scalable ... Define SLOs, improve observability, and lead incident response and root cause analysis * Platform ...

next page

Showing results 1-20

Senior Observability Engineer information

See Detroit, MI salary details

$58.9K

$125.3K

$181.7K

How much do senior observability engineer jobs pay per year?

As of Sep 7, 2026, the average yearly pay for senior observability engineer in Detroit, MI is $125,287.00, according to ZipRecruiter salary data. Most workers in this role earn between $103,500.00 and $142,100.00 per year, depending on experience, location, and employer.

What is a senior observability engineer?

A Senior Observability Engineer is a seasoned IT professional responsible for designing, implementing, and maintaining systems that monitor and provide insights into the performance, health, and reliability of software applications and infrastructure. They utilize tools for logging, monitoring, tracing, and alerting to ensure that systems are observable and any issues can be quickly detected and resolved. In addition to technical expertise, they often collaborate with development and operations teams to establish best practices, improve incident response, and optimize system performance. Their work is crucial for maintaining uptime, enhancing customer experiences, and supporting the scalability of technology platforms.

How does a senior observability engineer typically collaborate with development and operations teams?

A Senior Observability Engineer works closely with both development and operations teams to ensure robust monitoring, logging, and tracing solutions are in place across all applications and infrastructure. They often participate in architecture discussions to advise on best practices for instrumenting code and systems for observability. By analyzing metrics and alerting patterns, they help teams proactively resolve issues and optimize system performance. This role also involves mentoring engineers on observability tools and fostering a culture of transparency and accountability in incident response.

What are the key skills and qualifications needed to thrive as a senior observability engineer, and why are they important?

To thrive as a Senior Observability Engineer, you need expertise in monitoring, logging, and tracing systems, with a solid background in computer science or a related field. Familiarity with tools like Prometheus, Grafana, ELK stack, and cloud platforms, as well as certifications such as AWS Certified DevOps Engineer, are typically required. Strong problem-solving, collaboration, and communication skills are critical for effectively diagnosing and resolving complex infrastructure issues. These skills ensure reliable system performance, rapid incident response, and continuous improvement of the technology environment.

What is the difference between Senior Observability Engineer vs Site Reliability Engineer?

AspectSenior Observability EngineerSite Reliability Engineer
CredentialsExperience with monitoring tools, scripting, cloud platformsSame as Senior Observability Engineer, often with SRE certifications
Work EnvironmentFocus on monitoring, logging, and tracing systemsFocus on system reliability, automation, and incident response
Industry UsageUsed in tech companies emphasizing system observabilityCommon in large-scale tech and cloud services
Search/Comparison IntentOften compared for monitoring rolesCompared for reliability and system stability roles

While both roles require expertise in cloud platforms and scripting, the Senior Observability Engineer primarily focuses on designing and maintaining monitoring, logging, and tracing systems to ensure system visibility. In contrast, a Site Reliability Engineer emphasizes system reliability, automation, and incident management to maintain service uptime. Both roles are vital in tech environments but serve different core functions related to system health and stability.

How much do senior observability engineers make?

Senior observability engineers typically earn between $110,000 and $160,000 annually, depending on experience, location, and company size. They often work with tools like Prometheus, Grafana, and cloud platforms, and may require advanced knowledge of monitoring, logging, and alerting systems.

What does a senior observability engineer do?

A senior observability engineer designs, implements, and maintains systems to monitor the performance and health of software applications and infrastructure. They utilize tools like Prometheus, Grafana, and ELK stack to analyze metrics, logs, and traces, ensuring system reliability and performance. This role often requires strong scripting skills and knowledge of cloud environments and distributed systems.

What are popular job titles related to Senior Observability Engineer jobs in Detroit, MI?

For Senior Observability Engineer jobs in Detroit, MI, the most frequently searched job titles are:

What job categories do people searching Senior Observability Engineer jobs in Detroit, MI look for?

The top searched job categories for Senior Observability Engineer jobs in Detroit, MI are:

Infographic showing various Senior Observability Engineer job openings in Detroit, MI as of August 2026, with employment types broken down into 84% Full Time, and 16% Contract. Highlights an 52% In-person, 6% Hybrid, and 42% Remote job distribution, with an average salary of $125,287 per year, or $60.2 per hour.

AI Systems Engineer - DevOps& Observability - Senior

Ernst & Young Oman

Detroit, MI โ€ข On-site

$107 - $177/hr

Other

Medical, Dental, Retirement, PTO

Posted 2 days ago

New


Job description

Location: Anywhere in Country

At EY, weโ€™re all in to shape your future with confidence.

Weโ€™ll help you succeed in a globally connected powerhouse of diverse teams and take your career wherever you want it to go. Join EY and help to build a better working world.

The opportunity

We are seeking an AI Systems Engineer to own the delivery, model-serving, routing, and observability layer of EYโ€™s AI-native platform. These are the systems that ship, run, and make fully visible every AI workload. Within the Hybrid AI Multi-Environment Runtime (HAI), this role advised how AI services and agents are built and deployed, how models execute, how requests are routed to them, how AI assets are catalogued and governed, how consumption is measured and bounded, and how the entire platform is observed across cloud, onโ€‘prem, edge, and airโ€‘gapped environments. Works with senior engineers to test and develop capabilities.

This is a distinct discipline from platform, data, and trust engineering. Where Platform Engineering owns the cluster substrate and its infrastructure automation, this role owns the delivery and runtime surface, including the CI/CD/CV pipelines that ship AI workloads, secure model execution, semantic routing, and model/prompt selection, together with the governance, discovery, cost, and telemetry systems that keep AI workloads shippable, economical, discoverable, and transparent. It sits at the intersection of DevOps, MLOps, FinOps, and observability.

This role is ideal for an engineer who is equally comfortable building automated delivery pipelines, operating highโ€‘performance inference (GPUs, model servers, sandboxed execution), and building deep observability and cost visibility; who understands that in regulated contexts every AI workload must be delivered repeatably and every AI request must be economically bounded, attributable, and traceable endโ€‘toโ€‘end.

Your key responsibilities
  • Supports DevOps and delivery for AI workloads: build and operate the CI/CD/CV pipelines that ship AI services, agents, and runtime components, including automated build, test, continuous verification, release, and rollback, so AI workloads are delivered repeatably and safely into every environment.

  • Own governance and discovery for AI assets, including service catalog/registry (Artifactory/Nexus, Harbor), experiment tracking and model metadata (MLflow), upstream registries/mirrors (HuggingFace/NGC), CVE/SBOM scanning (Trivy), lineage contracts (OpenLineage), and license management.

  • Own resource and cost management, including quotas and rate limits, cost attribution and utilization (Apptio/OpenCost/Kubecost), so AI execution stays economically bounded and controllable per tenant and engagement.

  • Own the full observability stack, including metrics (Prometheus/Mimir), logs (Loki), traces (Tempo/Jaeger), dashboards (Grafana), LLM debugging and evaluation (LangSmith/Langfuse), and SLA/alert notifications.

  • Own the OpenTelemetry collection layer, including multiโ€‘tenant receiver, exporters and queues (Kafka sink), DCGM exporter for GPU telemetry, processor batching, and dynamic filtering, so every signal is captured and routed reliably.

  • Automate GitOpsโ€‘based delivery and continuous verification; embedding quality, integrity, and cost gates into pipelines so releases are policyโ€‘compliant by default rather than by manual review.

  • Close the loop between delivery and observability by using telemetry, evaluation, and cost signals to drive deployment decisions, progressive rollout, and automated rollback of AI workloads.

  • Ensure cost and telemetry are identityโ€‘stamped and perโ€‘tenant, so consumption and behavior are attributable endโ€‘toโ€‘end, keeping FinOps and observability tied to the workloads that generate the load.

Skills and attributes for success
  • Strong DevOps expertise: CI/CD/CV pipeline design, GitOps, continuous verification, and progressive/automated release and rollback for production workloads.

  • Deep expertise operating modelโ€‘serving and inference systems (Ray, vLLM/Triton/NIM) on GPUs at production scale.

  • Deep observability skills: metrics, logs, traces, and OpenTelemetry.

  • FinOps mindset: able to attribute, bound, and optimize AI consumption cost per tenant and workload.

  • Familiarity with model/artifact governance, registries, CVE scanning, and license/lineage tracking.

  • Comfortable operating across cloud, onโ€‘prem, edge, and airโ€‘gapped environments with consistent runtime and telemetry semantics.

  • Strong communicator able to explain runtime, cost, and observability tradeoffs to engineers, architects, and leadership.

To qualify you must have
  • 8+ years in DevOps, MLOps, platform, or observability engineering, with handsโ€‘on production ownership of AI or highโ€‘throughput services.

  • Strong handsโ€‘on DevOps experience, including CI/CD/CV pipelines and GitOps tooling (ArgoCD, Helm, GitHub Actions/GitLab CI, or equivalents) for automated build, test, release, and rollback.

  • Handsโ€‘on expertise operating inference/modelโ€‘serving frameworks (Ray Serve, vLLM, Triton, or NIM) on GPU infrastructure.

  • Strong experience with observability stacks (Prometheus, Grafana, Loki, Tempo/Jaeger) and OpenTelemetry.

  • Experience with API gateways and request routing (Envoy or equivalent), including streaming responses.

  • Experience with cost management / FinOps tooling (OpenCost, Kubecost, or equivalent) and quota/rateโ€‘limit enforcement.

  • Familiarity with model/artifact registries and supplyโ€‘chain scanning (Harbor, MLflow, Trivy/SBOM).

  • Proven track record operating AI or service infrastructure under compliance, security, or regulatory constraints.

  • Ability to define clean ownership boundaries and consumption contracts with platform, trust, and data teams.

Ideally, youโ€™ll also have
  • Bachelorโ€™s or Masterโ€™s degree in Computer Science or related technical field.

  • Experience with LLM evaluation and debugging tooling (LangSmith, Langfuse) and prompt/response quality measurement.

  • Experience with sandboxed/secure execution (gVisor, Firecracker, or microVM isolation) for untrusted or multiโ€‘tenant workloads.

  • Familiarity with GPU telemetry (DCGM) and GPU utilization optimization.

  • Experience with lineage and governance contracts (OpenLineage) and AI license management.

  • Exposure to multiโ€‘tenant cost attribution and perโ€‘tenant SLA/alerting.

  • Exposure to regulated delivery environments (financial services, tax, healthcare, risk).

What we offer you

At EY, weโ€™ll develop you with futureโ€‘focused skills and equip you with worldโ€‘class experiences. Weโ€™ll empower you in a flexible environment, and fuel you and your extraordinary talents in a diverse and inclusive culture of globally connected teams. Learn more .

  • We offer a comprehensive compensation and benefits package where youโ€™ll be rewarded based on your performance and recognized for the value you bring to the business. The base salary range for this job in all geographic locations in the US is $106,900 to $176,500. The base salary range for New York City Metro Area, Washington State and California (excluding Sacramento) is $128,400 to $200,600. Individual salaries within those ranges are determined through a wide variety of factors including but not limited to education, experience, knowledge, skills and geography. In addition, our Total Rewards package includes medical and dental coverage, pension and 401(k) plans, and a wide range of paid time off options.

  • Join us in our teamโ€‘led and leaderโ€‘enabled hybrid model. Our expectation is for most people in external, client serving roles to work together in person 40-60% of the time over the course of an engagement, project or year.

  • Under our flexible vacation policy, youโ€™ll decide how much vacation time you need based on your own personal circumstances. Youโ€™ll also be granted time off for designated EY Paid Holidays, Winter/Summer breaks, Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional wellโ€‘being.

EY focuses on highโ€‘ethical standards and integrity among its employees and expects all candidates to demonstrate these qualities.

EY | Building a better working world

EY is building a better working world by creating new value for clients, people, society and the planet, while building trust in capital markets.

Enabled by data, AI and advanced technology, EY teams help clients shape the future with confidence and develop answers for the most pressing issues of today and tomorrow.

EY teams work across a full spectrum of services in assurance, consulting, tax, strategy and transactions. Fueled by sector insights, a globally connected, multiโ€‘disciplinary network and diverse ecosystem partners, EY teams can provide services in more than 150 countries and territories.

EY provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, pregnancy, genetic information, national origin, protected veteran status, disability status, or any other legally protected basis, including arrest and conviction records, in accordance with applicable law.

EY is committed to providing reasonable accommodation to qualified individuals with disabilities including veterans with disabilities. If you have a disability and either need assistance applying online or need to request an accommodation during any part of the application process, please call 1-800-EY-HELP3, select Option 2 for candidate related inquiries, then select Option 1 for candidate queries and finally select Option 2 for candidates with an inquiry which will route you to EYโ€™s Talent Shared Services Team (TSS) or email the TSS at ssc.customersupport@ey.com .

#J-18808-Ljbffr