Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ... Manage and evolve Helm chart definitions and ArgoCD GitOps workflows for multi-region SaaS ...
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ... Manage and evolve Helm chart definitions and ArgoCD GitOps workflows for multi-region SaaS ...
... and Reliability Engineering team, responsible for ensuring the stability, availability, and ... Manage and support OpenShift platform operations, including container deployments, pod ...
... and Reliability Engineering team, responsible for ensuring the stability, availability, and ... Manage and support OpenShift platform operations, including container deployments, pod ...
Senior Site Reliability Engineer
Ottawa, ON · On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ... Manage and evolve Helm chart definitions and ArgoCD GitOps workflows for multi-region SaaS ...
Senior Site Reliability Engineer
Ottawa, ON · On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ... Manage and evolve Helm chart definitions and ArgoCD GitOps workflows for multi-region SaaS ...
Site Reliability Engineer
Kitchener, ON · On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Site Reliability Engineer
Kitchener, ON · On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Corporate Reliability Engineering Co-Op
CA$22 - CA$25/hr
... Management System (CMMS)with interpretation and trending of highly critical equipment with common repeating failure patterns at the plant * Supporting Reliability Engineer Data collection and ...
Corporate Reliability Engineering Co-Op
CA$22 - CA$25/hr
... Management System (CMMS)with interpretation and trending of highly critical equipment with common repeating failure patterns at the plant * Supporting Reliability Engineer Data collection and ...
Senior Site Reliability Engineer
Waterloo, ON · On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ... Manage and evolve Helm chart definitions and ArgoCD GitOps workflows for multi-region SaaS ...
Senior Site Reliability Engineer
Waterloo, ON · On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ... Manage and evolve Helm chart definitions and ArgoCD GitOps workflows for multi-region SaaS ...
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON · Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
Quick apply
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON · Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
Site Reliability Engineer
Mississauga, ON · Hybrid
We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ... Create, manage, and support pipelines that the application support teams will be utilizing to ...
Site Reliability Engineer
Mississauga, ON · Hybrid
We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ... Create, manage, and support pipelines that the application support teams will be utilizing to ...
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON · Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service ... Reporting to the Sr Manager, Applications Operations Engineering Key Responsibilities: Participate ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON · Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service ... Reporting to the Sr Manager, Applications Operations Engineering Key Responsibilities: Participate ...
Role Summary The Lead, Site Reliability Engineering ensures monitoring and analysis is conducted to ... manage the software development process. The Lead, DEV Platform Support Engineer will play a ...
Role Summary The Lead, Site Reliability Engineering ensures monitoring and analysis is conducted to ... manage the software development process. The Lead, DEV Platform Support Engineer will play a ...
We're searching for a Senior Site Reliability Engineer who not only brings deep technical expertise ... Create and refine tooling and processes that improve incident management, driving proactive ...
We're searching for a Senior Site Reliability Engineer who not only brings deep technical expertise ... Create and refine tooling and processes that improve incident management, driving proactive ...
Senior Site Reliability Engineer - Azure
Toronto, ON · Hybrid
CA$99K/yr
... Manage infrastructure deployment pipelines and troubleshoot onboarding and operational issues ... in Site Reliability, DevOps, or Cloud Engineering roles Expertise with Microsoft Azure; AWS ...
Senior Site Reliability Engineer - Azure
Toronto, ON · Hybrid
CA$99K/yr
... Manage infrastructure deployment pipelines and troubleshoot onboarding and operational issues ... in Site Reliability, DevOps, or Cloud Engineering roles Expertise with Microsoft Azure; AWS ...
Technical and Management Escalation point for Service Operations Centre (SOC) engineers and during ... Experience working in a similar site reliability role This role offers great perks and a ...
Quick apply
Technical and Management Escalation point for Service Operations Centre (SOC) engineers and during ... Experience working in a similar site reliability role This role offers great perks and a ...
Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is ... Participate in the on-call rotation, and incident management Required Knowledge, Skills, and ...
Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is ... Participate in the on-call rotation, and incident management Required Knowledge, Skills, and ...
At its heart, the Smile platform enables people and organizations to better manage healthcare data ... The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability ...
At its heart, the Smile platform enables people and organizations to better manage healthcare data ... The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability ...
Technical and Management Escalation point for Service Operations Centre (SOC) engineers and during ... Experience working in a similar site reliability role This role offers great perks and a ...
Quick apply
Technical and Management Escalation point for Service Operations Centre (SOC) engineers and during ... Experience working in a similar site reliability role This role offers great perks and a ...
Technical and Management Escalation point for Service Operations Centre (SOC) engineers and during ... Experience working in a similar site reliability role This role offers great perks and a ...
Quick apply
Technical and Management Escalation point for Service Operations Centre (SOC) engineers and during ... Experience working in a similar site reliability role This role offers great perks and a ...
Technical and Management Escalation point for Service Operations Centre (SOC) engineers and during ... Experience working in a similar site reliability role This role offers great perks and a ...
Quick apply
Technical and Management Escalation point for Service Operations Centre (SOC) engineers and during ... Experience working in a similar site reliability role This role offers great perks and a ...
Lead Site Reliability Administrator
Mississauga, ON · On-site
CA$177K/yr
As a Site Reliability Engineer (SRE) at OpenText, you will be responsible for ensuring the ... Architect and manage Kubernetes clusters supporting pods across multi-region deployments to ensure ...
New
Lead Site Reliability Administrator
Mississauga, ON · On-site
CA$177K/yr
As a Site Reliability Engineer (SRE) at OpenText, you will be responsible for ensuring the ... Architect and manage Kubernetes clusters supporting pods across multi-region deployments to ensure ...
New
Reliability Engineer Manager information
What does a reliability engineer manager do?
What are some common challenges reliability engineer managers face when balancing long-term reliability improvements with immediate operational demands?
What are the key skills and qualifications needed to thrive as a reliability engineer manager?
What is the difference between Reliability Engineer Manager vs Reliability Engineer?
| Aspect | Reliability Engineer | Reliability Engineer Manager |
|---|---|---|
| Required Credentials | Bachelor's in Engineering or related field; certifications like CRC, CRE | Same as Reliability Engineer, plus leadership experience |
| Work Environment | Design, analyze, and improve system reliability; often in teams | Oversees Reliability Engineers; manages projects and teams |
| Employer & Industry Usage | Manufacturing, aerospace, energy, automotive | Same industries, with added managerial responsibilities |
| Common Search & Comparison | Focuses on technical skills and hands-on reliability tasks | Focuses on leadership, team management, and strategic planning |
The main difference between a Reliability Engineer and a Reliability Engineer Manager lies in their responsibilities. The Reliability Engineer focuses on technical analysis and system improvements, while the Reliability Engineer Manager oversees teams, manages projects, and develops strategies to enhance reliability across the organization.
What are the most commonly searched types of Reliability Engineer jobs in Ontario?
The most popular types of Reliability Engineer jobs in Ontario are:
What are popular job titles related to Reliability Engineer Manager jobs in Ontario?
For Reliability Engineer Manager jobs in Ontario, the most frequently searched job titles are:
What job categories do people searching Reliability Engineer Manager jobs in Ontario look for?
The top searched job categories for Reliability Engineer Manager jobs in Ontario are:
What cities in Ontario are hiring for Reliability Engineer Manager jobs?
Cities in Ontario with the most Reliability Engineer Manager job openings:

Full-time
Medical, Retirement
Re-posted 15 days ago
Job description
- Own and operate production Kubernetes clusters (Amazon EKS) including upgrades, scaling, security hardening, and cluster lifecycle management;
Design, implement, and maintain infrastructure-as-code using Terraform; contribute to shared module libraries and enforce IaC standards across the team;
Manage and evolve Helm chart definitions and ArgoCD GitOps workflows for multi-region SaaS deployments;
Operate and maintain observability infrastructure including Grafana, alerts, dashboards, and log pipelines. Act to eliminate noise and surface signal;
Contribute to pipeline reliability: identify flaky stages, reduce build times, improve developer experience across CI/CD pipelines;
Remediate security vulnerabilities (CVEs) in container images and infrastructure components; participate in compliance work including FedRAMP support activities;
Develop and maintain runbooks, change management procedures, and operational documentation;
Ensure alignment with internal policies and frameworks such as ISO 27001, SOC2, and NIST;
Contribute to AI-assisted tooling and automation (e.g., Claude-based Terraform agents, automated triage tools) as part of the team's operational efficiency roadmap;
Participate in on-call incident response rotation; lead or support incident command during active production incidents including root cause analysis and post-incident review.
5+ years of industry experience with a trajectory that demonstrates growing depth in cloud infrastructure and SRE practices;
Managed production Kubernetes environments at scale: not just deployed workloads, but owned cluster health, upgrades, and failure modes;
Responded to production incidents in high-stakes environments where downtime has real consequences;
Written and maintained Terraform at the module level, not just as a consumer: understands state, dependencies, and the operational burden of drift;
Operated in an environment that uses GitOps: has a good understanding of Helm chart organization, ArgoCD app-of-apps patterns, or equivalent;
Balanced reactive operational work with proactive roadmap delivery; knows how to protect time for improvements while keeping production stable;
Worked with observability as a first-class discipline: built meaningful dashboards, eliminated alert fatigue, and used metrics to make operational decisions;
Contributed to security hardening in a regulated or compliance-adjacent environment: FedRAMP, SOC 2, or similar frameworks are a strong asset.