1

Observability Manager Jobs in Toronto, ON (NOW HIRING)

Manage and optimize CiscoThousandEyesfor network path visibility, internet performance monitoring, and end-user experience insights. * Define and enforce observability-as-code practices, managing ...

Manage and optimize CiscoThousandEyesfor network path visibility, internet performance monitoring, and end-user experience insights. * Define and enforce observability-as-code practices, managing ...

Manager, Valuations - Commodities

Toronto, ON · On-site

CA$96K - CA$136K/yr

In-depth understanding of fair value measurement, market observability, valuation adjustments ... Ability to manage multiple deliverables in a fast-paced, evolving environment, including month-end ...

Build Test Data and Observability Systems Drive the creation of scalable, privacy & regulatory compliant test data management strategies. Integrate logs, traces, and monitoring into automated test ...

Observability Solutions play a critical role within the bank to provide monitoring coverage to all ... Understands importance of source code configuration management * Tracks open issues and resolution ...

New

next page

Showing results 1-20

Observability Manager information

What is the difference between Observability Manager vs Site Reliability Engineer?

AspectObservability ManagerSite Reliability Engineer
CredentialsTypically requires experience in monitoring, logging, and cloud tools; certifications like AWS, Google Cloud, or Kubernetes are commonRequires strong background in systems engineering, scripting, and cloud platforms; certifications like AWS, GCP, or Linux are often preferred
Work EnvironmentFocuses on overseeing observability tools, data analysis, and team coordination in tech environmentsHands-on role involving system automation, incident response, and infrastructure reliability
Industry UsageUsed across tech companies to improve system visibility and performanceCommon in DevOps and SRE teams to ensure system reliability and uptime

The Observability Manager primarily oversees monitoring and logging strategies, ensuring system visibility, while the Site Reliability Engineer is more hands-on, focusing on automating infrastructure and maintaining system reliability. Both roles require technical expertise and often collaborate closely but differ in scope and daily responsibilities.

What are the most commonly searched types of Observability jobs in Toronto, ON? The most popular types of Observability jobs in Toronto, ON are:
What job categories do people searching Observability Manager jobs in Toronto, ON look for? The top searched job categories for Observability Manager jobs in Toronto, ON are:
Infographic showing various Observability Manager job openings in Toronto, ON as of August 2026, with employment types broken down into 100% Full Time. Highlights an 67% In-person, and 33% Remote job distribution.

Sr. Observability Engineering

McKesson

Mississauga, ON • On-site, Remote

Full-time

Posted 24 days ago


McKesson rating

7.9

Company rating: 7.9 out of 10

Based on 209 frontline employees who took The Breakroom Quiz

48th of 86 rated pharmaceutical


Job description

McKesson is an impact-driven, Fortune 10 company that touches virtually every aspect of healthcare. We are known for delivering insights, products, and services that make quality care more accessible and affordable. Here, we focus on the health, happiness, and well-being of you and those we serve - we care.

What you do at McKesson matters. We foster a culture where you can grow, make an impact, and are empowered to bring new ideas. Together, we thrive as we shape the future of health for patients, our communities, and our people. If you want to be part of tomorrow's health today, we want to hear from you.

POSITION SUMMARY

McKesson is seeking a Senior Observability Engineer to join our Platform Engineering team. In this role, you will serve as a keyengineerand practitioner of McKesson's enterprise observability strategy,supportingthe design, deployment, and continuous improvement of our monitoring and telemetry platforms across cloud and on-premises infrastructure. You will operate with a high degree of autonomy, applying SRE principles and an engineering-first mindset to ensure the reliability, performance, and availability of mission-critical healthcare technology systems.

You will influence observability standards across the organization, mentor colleagues, and drive innovation without direct management responsibilities. Promotions beyond this level are selective and driven by demonstrated business impact.

KEY RESPONSIBILITIES

Observability Platform Ownership

  • Engineer, implement, and operate enterprise-grade observability platforms including Dynatrace,LogicMonitor, Grafana, and Prometheus across multi-cloud and hybrid environments.

  • Lead the adoption and integration ofOpenTelemetry(OTel) standards for distributed tracing, metrics collection, and log correlation across engineering teams.

  • Manage and optimize CiscoThousandEyesfor network path visibility, internet performance monitoring, and end-user experience insights.

  • Define and enforce observability-as-code practices, managing configurations through version-controlled pipelines (e.g., Terraform, Helm,GitOps).

  • Evaluate, recommend, and pilot emerging observability technologies and methodologies to continuously advance McKesson's monitoring capabilities.

SRE & Engineering Excellence

  • Apply Site Reliability Engineering (SRE) practices including error budgets, toil reduction, and chaos engineering to enhance system resilience.

  • Define, implement, and report on Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Key Performance Indicators (KPIs) in partnership with engineering and business stakeholders.

  • Develop and maintain alerting frameworks that minimize noise and maximize signal fidelity, reducing mean time to detect (MTTD) and mean time to resolve (MTTR).

  • Build self-service observability tooling and dashboards that enable engineering teams to own their own reliability metrics.

  • Lead blameless post-incident reviews (PIRs) and drive remediation actions to prevent recurrence.

Cloud & Infrastructure Monitoring

  • Design comprehensiveobservability strategies for cloud-native workloads across AWS, Azure, and/or GCP, including containers (Kubernetes/EKS/AKS), serverless, and microservices architectures.

  • Instrument infrastructure, application, and business-layer telemetry (logs, metrics, traces) to provide end-to-end visibility.

  • Manage network performance monitoring and synthetic testing viaThousandEyesto proactively identify and resolve connectivity and latency issues impacting end users.

  • Partner with DevOps, NetOps, and Security teams to integrate observability signals into CI/CD pipelines, ITSM workflows, and SOC operations.

AI & LLM Observability (Emerging Capability)

  • Extend the organization's existing observability platforms to instrument AI-powered applications and autonomous agent workflows, capturing the telemetry unique to non-deterministic systems.

  • ImplementOpenTelemetryGenAI semantic conventions to standardize how AI agent telemetry (traces, spans, token usage, tool invocations) is collected across frameworks and services.

  • Design and maintain tracing pipelines that capture the full agent execution chain from user intent and planner decisions through tool calls, retrieval steps, and model responses enabling root cause analysis beyond traditional request/response tracing.

  • Define AI-specific SLOs and cost budgets covering token consumption, hallucination rate, tool invocation success rate, and agent task completion rate.

  • Build dashboards and alerting for AI behavioral signals: model drift, performance degradation, prompt/context version changes, and guardrail violations.

  • Collaborate with data science,MLOps, and application teams to integrate observability signals into AI evaluation pipelines and continuous improvement feedback loops.

  • Implement security-focused AI monitoring including prompt injection detection, PII exposure in logs, and anomalous tool call patterns in alignment with OWASP Top 10 for LLMs.

Collaboration & Technical Leadership

  • Act as an observability subject-matter expert and trusted resource for engineers, architects, and product teams across McKesson.

  • Develop and maintain internal observability standards, runbooks, and best-practice documentation.

  • Mentor and guide junior and mid-level engineers on observability tooling, instrumentation techniques, and SRE principles.

  • Anticipate organizational scaling needs and proactively direct engineering efforts to stay ahead of capacity, reliability, and visibility challenges.

  • Contribute to the development of new frameworks, methodologies, and tooling that advance McKesson's engineering maturity.

REQUIRED QUALIFICATIONS

  • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field, or equivalent practical experience.

  • 7+ years of experience in observability, monitoring, SRE, or platform engineering roles within complex, enterprise-scale environments.

  • Hands-on expertise with two or more of the following platforms: Dynatrace,LogicMonitor,Grafana,ThousandEyesandPrometheus.

  • Demonstrated experience implementingOpenTelemetry(OTel) for distributed tracing and metrics instrumentation in polyglot service environments.

  • Proficiency with CiscoThousandEyesor similar platformsfor network monitoring, synthetic testing, and internet intelligence.

  • Deep understanding of cloud infrastructure monitoring across AWS, Azure, and/or GCP, including IaaS, PaaS, and container platforms.

  • Strong working knowledge of SRE principles: SLOs, SLAs, error budgets, alerting philosophy, and incident management.

  • Scripting/automation proficiency in Python, Go, Bash, or equivalent, with experience building or extending monitoring integrations.

  • Experience with infrastructure-as-code andGitOpspractices (Terraform, Helm, Ansible, etc.) as applied to observability configuration management.

  • Demonstrated ability to work independently, set direction, and deliver results in a complex, matrixed organization.

PREFERRED QUALIFICATIONS

  • Experience in healthcare IT, regulated industries, or large-scale enterprise environments.

  • Familiarity with AIOps platforms and ML-based anomaly detection capabilities within Dynatrace or similar tooling.

  • Knowledge ofeBPF-based observability approaches and agent-less instrumentation techniques.

  • Experience integrating observability data into ITSM platforms (ServiceNow, PagerDuty) and security workflows (SIEM).

Relocation is not budgeted for this role

Office requirement: We are Flex and Connect with 2 days a week in office.

We are proud to offer a competitive compensation package at McKesson as part of our Total Rewards. This is determined by several factors, including performance, experience and skills, equity, regular job market evaluations, and geographical markets. The pay range shown below is aligned with McKesson's pay philosophy, and pay will always be compliant with any applicable regulations. In addition to base pay, other compensation, such as an annual bonus or long-term incentive opportunities may be offered. For more information regarding benefits at McKesson, pleaseclick here.

Our Base Pay Range for this position

$99,100 - $132,100

McKesson has become aware of online recruiting-related scams in which individuals who are not affiliated with or authorized by McKesson are using McKesson's (or affiliated entities, like CoverMyMeds or RxCrossroads) name in fraudulent emails, job postings or social media messages. In light of these scams, please bear the following in mind:
McKesson Talent Advisors will never solicit money or credit card information in connection with a McKesson job application.


McKesson Talent Advisors do not communicate with candidates via online chatrooms or using email accounts such as Gmail or Hotmail. Note that McKesson does rely on a virtual assistant (Gia) for certain recruiting-related communications with candidates.

McKesson job postings are posted on our career site: careers.mckesson.com.

McKesson is an Equal Opportunity Employer

McKesson provides equal employment opportunities to applicants and employees, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, age, genetic information, or any other legally protected category. For additional information on McKesson's full Equal Employment Opportunity policies, visit our Equal Employment Opportunity page.

McKesson is committed to being an Equal Employment Opportunity Employer and offers opportunities to all job seekers including job seekers with disabilities. If you need a reasonable accommodation to assist with your job search or application for employment, please contact us by sending an email to (United States) Disability_Accommodation@McKesson.com or (Canada) Accessibility@mckesson.ca. Resumes or CVs submitted to this email box will not be accepted.

Join us at McKesson!


What McKesson employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom