Enhance platform observability through metrics, logs, tracing, and actionable alerting to improve ... Mentor and guide engineers on cloud-native technologies, site reliability engineering principles ...
Enhance platform observability through metrics, logs, tracing, and actionable alerting to improve ... Mentor and guide engineers on cloud-native technologies, site reliability engineering principles ...
Manager, Site Reliability Engineering and DevOps
Toronto, ON · Hybrid
CA$142K - CA$186K/yr
... SRE engineers supporting production systems across cloud and onpremises environments, driving service reliability through Service Level Objectives (SLOs), observability, and automation. You will ...
Manager, Site Reliability Engineering and DevOps
Toronto, ON · Hybrid
CA$142K - CA$186K/yr
... SRE engineers supporting production systems across cloud and onpremises environments, driving service reliability through Service Level Objectives (SLOs), observability, and automation. You will ...
Site Reliability Engineer
Toronto, ON · Hybrid
Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring the ... Observability and monitoring tools Soft Skills: * Strong cross-functional collaboration and ...
Site Reliability Engineer
Toronto, ON · Hybrid
Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring the ... Observability and monitoring tools Soft Skills: * Strong cross-functional collaboration and ...
SRE is part of a global organization that leverages the latest technology to communicate with our ... You'll be a key voice in observability, change management, and service scalability, providing ...
Quick apply
SRE is part of a global organization that leverages the latest technology to communicate with our ... You'll be a key voice in observability, change management, and service scalability, providing ...
Lead the transformation toward a modern SRE-driven operating model incorporating AIOps and NoOps principles * Drive implementation of observability frameworks, predictive analytics, and event ...
Lead the transformation toward a modern SRE-driven operating model incorporating AIOps and NoOps principles * Drive implementation of observability frameworks, predictive analytics, and event ...
As a Consultant, Site Reliability Engineering, you'll apply advanced software engineering ... Observability and Telemetry - Define, implement, and maintain observability strategies, including ...
As a Consultant, Site Reliability Engineering, you'll apply advanced software engineering ... Observability and Telemetry - Define, implement, and maintain observability strategies, including ...
As a Consultant, Site Reliability Engineering, you'll apply advanced software engineering ... Observability and Telemetry - Define, implement, and maintain observability strategies, including ...
As a Consultant, Site Reliability Engineering, you'll apply advanced software engineering ... Observability and Telemetry - Define, implement, and maintain observability strategies, including ...
Senior SRE/AIOps Engineer
Toronto, ON · On-site
Build dashboards and observability KPIs for operational insights Automation & Self-Healing Systems ... Must Have: * 3+ years of SRE or Systems Engineering experience with strong technical expertise.
Senior SRE/AIOps Engineer
Toronto, ON · On-site
Build dashboards and observability KPIs for operational insights Automation & Self-Healing Systems ... Must Have: * 3+ years of SRE or Systems Engineering experience with strong technical expertise.
Drive continuous improvement of observability maturity, service visibility, monitoring ... Act as the Compute SRE representative for enterprise monitoring strategy, operational intelligence ...
Drive continuous improvement of observability maturity, service visibility, monitoring ... Act as the Compute SRE representative for enterprise monitoring strategy, operational intelligence ...
The successful candidate will apply SRE practices such as incident response, observability, automation, runbook development, root-cause analysis, and continuous reliability improvement while ...
The successful candidate will apply SRE practices such as incident response, observability, automation, runbook development, root-cause analysis, and continuous reliability improvement while ...
As Head of SRE and Production Operations, you will lead the vision, design, development ... Enable automated end-to-end observability across all critical applications within Commercial ...
As Head of SRE and Production Operations, you will lead the vision, design, development ... Enable automated end-to-end observability across all critical applications within Commercial ...
We treat reliability as a product - prioritizing observability, automation, and safe operations so ... AI SRE exists to ensure PointClickCare's AI platforms run safely, reliably, and efficiently ...
We treat reliability as a product - prioritizing observability, automation, and safe operations so ... AI SRE exists to ensure PointClickCare's AI platforms run safely, reliably, and efficiently ...
Site Reliability Engineer
Toronto, ON · On-site +1
CA$125K - CA$250K/yr
We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind ... Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems
Site Reliability Engineer
Toronto, ON · On-site +1
CA$125K - CA$250K/yr
We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind ... Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems
Site Reliability Engineer - Cloud & Platform Engineering, Manulife Bank Technology
Toronto, ON · On-site
Join Manulife Bank Technology team as a Site Reliability Engineer (SRE) and help deliver reliable ... Develop and enhance observability capabilities, including metrics, logging, tracing, dashboards ...
Site Reliability Engineer - Cloud & Platform Engineering, Manulife Bank Technology
Toronto, ON · On-site
Join Manulife Bank Technology team as a Site Reliability Engineer (SRE) and help deliver reliable ... Develop and enhance observability capabilities, including metrics, logging, tracing, dashboards ...
Senior Site Reliability Developer
Toronto, ON · On-site
CA$107K - CA$157K/yr
... (SRE) to manage critical cloud infrastructure and site reliability operations for the Autodesk ... Knowledge of standardized observability frameworks such as OpenTelemetry * Relevant certifications ...
Senior Site Reliability Developer
Toronto, ON · On-site
CA$107K - CA$157K/yr
... (SRE) to manage critical cloud infrastructure and site reliability operations for the Autodesk ... Knowledge of standardized observability frameworks such as OpenTelemetry * Relevant certifications ...
Site Reliability Engineer Hybrid GTA 12+ months This is an opportunity for someone early in their career who has hands-on experience with Azure data technologies and wants to build deeper experience ...
Site Reliability Engineer Hybrid GTA 12+ months This is an opportunity for someone early in their career who has hands-on experience with Azure data technologies and wants to build deeper experience ...
Site Reliability Engineer
Toronto, ON · Hybrid
CA$100K - CA$125K/yr
As a Site Reliability Engineer, you will play a crucial role in enhancing the reliability, performance, and scalability of our systems and services. You will be a part of a global "commando" team of ...
Site Reliability Engineer
Toronto, ON · Hybrid
CA$100K - CA$125K/yr
As a Site Reliability Engineer, you will play a crucial role in enhancing the reliability, performance, and scalability of our systems and services. You will be a part of a global "commando" team of ...
Senior Site Reliability Engineer I
CA$153K - CA$277K/yr
Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing services ... Observability Systems: Experience using monitoring, profiling, and observability platforms (e.g ...
Senior Site Reliability Engineer I
CA$153K - CA$277K/yr
Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing services ... Observability Systems: Experience using monitoring, profiling, and observability platforms (e.g ...
Build and evolve observability, monitoring, and alerting pipelines (metrics, logs, traces) using ... Must Have: * 5+ years of experience in Site Reliability Engineering, DevOps or Platform Engineering ...
Build and evolve observability, monitoring, and alerting pipelines (metrics, logs, traces) using ... Must Have: * 5+ years of experience in Site Reliability Engineering, DevOps or Platform Engineering ...
As a Site Reliability Engineer (SRE) at OpenText, you will be responsible for ensuring the ... Azure DevOps certification or equivalent certifications. * Familiarity with observability tools ...
As a Site Reliability Engineer (SRE) at OpenText, you will be responsible for ensuring the ... Azure DevOps certification or equivalent certifications. * Familiarity with observability tools ...
Observability Site Reliability Engineer information
What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?
| Aspect | Observability Site Reliability Engineer | Monitoring Engineer |
|---|---|---|
| Focus | Ensuring system reliability through observability, automation, and incident response | Implementing and managing monitoring tools and dashboards |
| Skills | Cloud platforms, scripting, incident management, observability tools | Monitoring tools, alerting systems, data analysis |
| Work Environment | DevOps teams, cloud infrastructure, large-scale systems | Operations teams, infrastructure monitoring |
While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.
What are popular job titles related to Observability Site Reliability Engineer jobs in Toronto, ON?
For Observability Site Reliability Engineer jobs in Toronto, ON, the most frequently searched job titles are:
- Site Reliability Engineer
- Site Reliability Engineer Remote
- Sre Internship
- Site Reliability Engineer Intern
- Entry Level Site Reliability Engineer
- Senior Site Reliability Engineer
- Remote Process Engineer
- Remote Site Reliability Engineer Intern
- Remote Reliability Engineer
- Full Time Site Reliability Engineer Remote
What job categories do people searching Observability Site Reliability Engineer jobs in Toronto, ON look for?
The top searched job categories for Observability Site Reliability Engineer jobs in Toronto, ON are:

Full-time
Medical, Retirement
Posted 14 days ago
Job description
This is a hands-on senior engineering role focused on improving production resilience, strengthening security, driving operational excellence, and enhancing the developer experience across the organization.
In this role, you will design, build, and evolve the foundational systems, tooling, and operational practices that enable engineering teams to ship secure, reliable, and scalable software with confidence. You will help establish reliability standards, define service level objectives (SLOs), improve observability, automate operational processes, and drive incident management and post-incident learning practices that strengthen platform stability over time.
Partnering closely with Engineering, Security, Platform, and Product teams, you will architect scalable distributed systems, optimize Kubernetes and AWS-based infrastructure, and build automated delivery pipelines that support rapid and safe software releases. You will play a key role in reducing operational toil, improving system performance, increasing platform reliability, and ensuring that our infrastructure can support continued business growth.
This is a full-time permanent position
This is an existing vacancy
Location: This is a remote location open to candidates legally authorized to work in Canada.
- Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes.
- Design, implement, and continuously improve deployment, release, and rollback strategies across complex distributed systems.
- Establish secure-by-default CI/CD pipelines with robust automation, governance, and policy-driven controls.
- Enhance platform observability through metrics, logs, tracing, and actionable alerting to improve system visibility and operational efficiency.
- Define, implement, and mature Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability standards across the organization.
- Lead response efforts for high-severity incidents, ensuring timely resolution, effective communication, and meaningful post-incident reviews that drive continuous improvement.
- Partner closely with engineering teams to strengthen platform standards, improve service resilience, optimize runtime performance, and embed reliability best practices.
- Mentor and guide engineers on cloud-native technologies, site reliability engineering principles, and operational excellence practices, fostering a culture of continuous learning and accountability.
- 8+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or related cloud-native engineering roles.
- Deep expertise in AWS services, including EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security best practices.
- Advanced experience operating and scaling production Kubernetes environments.
- Strong hands-on experience with Istio service mesh, including traffic management, security, observability, and resiliency.
- Proven expertise with Infrastructure as Code (IaC), preferably using AWS CDK.
- Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms.
- Strong troubleshooting, performance optimization, and incident management experience in distributed systems.
- Excellent communication, collaboration, and technical leadership skills.
- Experience designing and operating monitoring, logging, tracing, and alerting solutions for cloud-native platforms.
- Strong knowledge of AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling.
- Experience defining and operationalizing SLIs, SLOs, alerting strategies, runbooks, and reliability metrics.
- Proven ability to leverage observability data to improve service reliability, reduce incident impact, and optimize operational performance.
- Strong proficiency in TypeScript and Node.js for platform engineering, automation, and operational tooling.
- Experience building and maintaining scalable backend services, APIs, and event-driven systems.
- Deep understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking.
- Experience implementing zero-trust architectures, mTLS, and service-to-service security controls.
- Commitment to high-quality engineering practices, including automated testing, code reviews, and observability-driven development.
- Strong understanding of resilience engineering, including autoscaling, disruption management, failure testing, and safe deployment strategies.
- Experience with progressive delivery practices such as canary, blue/green, and feature-flag-based deployments.
- Experience working in regulated, compliance-driven, or security-sensitive SaaS environments.
- Familiarity with FinOps principles and cost optimization strategies for cloud platforms.
- Experience building internal developer platforms and self-service engineering tooling.
- Cloud-native certifications such as CKA, CKAD, CKS, KCSA, or KCNA.
- Kubestronaut certification or equivalent advanced Kubernetes expertise is highly regarded.
Salary Range:
The annual base salary for this position is between $140,000 CAD and $155,000 CAD per year.
This role is also eligible for discretionary bonus and/or commission, as well as other benefits. Actual pay within the listed range will be determined based on factors such as transferable skills, relevant experience, market conditions, and primary work location. The posted range is subject to change and may be updated periodically.