Reporting to the SRE, Developer Platform Engineering , the Lead, DEV Platform Support Engineer will play a critical role in supporting and enhancing our Azure-based developer platforms and internal ...
Reporting to the SRE, Developer Platform Engineering , the Lead, DEV Platform Support Engineer will play a critical role in supporting and enhancing our Azure-based developer platforms and internal ...
Corporate Reliability Engineering Co-Op
CA$22 - CA$25/hr
Supporting Reliability Engineer Data collection and standardization of Failure modes and the Preventive (PM) and Predictive (PdM) Maintenance tasks to maintain the assets * Accurately create ...
Corporate Reliability Engineering Co-Op
CA$22 - CA$25/hr
Supporting Reliability Engineer Data collection and standardization of Failure modes and the Preventive (PM) and Predictive (PdM) Maintenance tasks to maintain the assets * Accurately create ...
Reliability Engineering Specialist
Brantford, ON ยท On-site
CA$95K - CA$125K/yr
As a Reliability Engineering Specialist, you will play a key role in improving asset reliability, reducing downtime, and optimizing maintenance strategies across the site. Working closely with ...
Reliability Engineering Specialist
Brantford, ON ยท On-site
CA$95K - CA$125K/yr
As a Reliability Engineering Specialist, you will play a key role in improving asset reliability, reducing downtime, and optimizing maintenance strategies across the site. Working closely with ...
Head of Platform Engineering, Reliability & Control
Toronto, ON ยท Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
Head of Platform Engineering, Reliability & Control
Toronto, ON ยท Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
As an Application Support Engineer for our Quantitative Investment Services, you'll take ownership of the reliability and performance of mission-critical trading platforms. This is more than a ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
As an Application Support Engineer for our Quantitative Investment Services, you'll take ownership of the reliability and performance of mission-critical trading platforms. This is more than a ...
Site Reliability Engineer
Kitchener, ON ยท On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Posted today
Site Reliability Engineer
Kitchener, ON ยท On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Posted today
Site Reliability Engineer
Vaughan, ON ยท On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Site Reliability Engineer
Vaughan, ON ยท On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON ยท Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service during incidents, and minimise customer impact. When issues occur, the role focuses not only on ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON ยท Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service during incidents, and minimise customer impact. When issues occur, the role focuses not only on ...
Network Reliability Engineer
Paris, ON ยท On-site
We follow the Network Reliability Engineer (NRE) model: continuous observation, rapid capacity adaptation, and relentless automation so that network capabilities are delivered as APIs, not tickets.
Network Reliability Engineer
Paris, ON ยท On-site
We follow the Network Reliability Engineer (NRE) model: continuous observation, rapid capacity adaptation, and relentless automation so that network capabilities are delivered as APIs, not tickets.
Director, Reliability Engineering
Toronto, ON ยท On-site
Job Summary The Director, Reliability Engineering is responsible for implementing and leading strategies to maximize uptime, performance and lifespan of assets, and reducing likelihood of failure by ...
Director, Reliability Engineering
Toronto, ON ยท On-site
Job Summary The Director, Reliability Engineering is responsible for implementing and leading strategies to maximize uptime, performance and lifespan of assets, and reducing likelihood of failure by ...
Senior Site Reliability Engineer
Waterloo, ON ยท On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Senior Site Reliability Engineer
Waterloo, ON ยท On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON ยท Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
Quick apply
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON ยท Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
We're searching for a Senior Site Reliability Engineer who not only brings deep technical expertise but also thrives on engaging with customers, partners, and cross-functional teams. In this customer ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
We're searching for a Senior Site Reliability Engineer who not only brings deep technical expertise but also thrives on engaging with customers, partners, and cross-functional teams. In this customer ...
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Quick apply
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Quick apply
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Senior Site Reliability Engineer - Azure
Toronto, ON ยท Hybrid
CA$99K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Senior Site Reliability Engineer, you will be embedded within one of our Product Areas, taking ownership of specific responsibility domains where your experience ...
Senior Site Reliability Engineer - Azure
Toronto, ON ยท Hybrid
CA$99K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Senior Site Reliability Engineer, you will be embedded within one of our Product Areas, taking ownership of specific responsibility domains where your experience ...
Senior Platform Reliability Engineer
Waterloo, ON ยท On-site
CA$113K - CA$163K/yr
Senior Platform Reliability Engineer Join Manulife Global Wealth & Asset Management (GWAM) and help power critical cloud platforms that enable data, analytics, and AI across the organization. We're ...
Senior Platform Reliability Engineer
Waterloo, ON ยท On-site
CA$113K - CA$163K/yr
Senior Platform Reliability Engineer Join Manulife Global Wealth & Asset Management (GWAM) and help power critical cloud platforms that enable data, analytics, and AI across the organization. We're ...
Reliability Engineer information
What are some typical challenges reliability engineers face when implementing preventive maintenance strategies?
How much do reliability engineers get paid?
What are the key skills and qualifications needed to thrive as a reliability engineer, and why are they important?
What is the difference between Reliability Engineer vs Maintenance Engineer?
| Aspect | Reliability Engineer | Maintenance Engineer |
|---|---|---|
| Credentials | Typically requires engineering degree, certifications in reliability or asset management | Often requires engineering or technical diploma, certifications in maintenance or equipment repair |
| Work Environment | Focuses on analysis, design, and improvement of systems for reliability | Hands-on maintenance, repair, and troubleshooting of equipment |
| Industry Usage | Common in manufacturing, energy, aerospace, and industrial sectors | Prevalent in manufacturing, facilities, and industrial plants |
Reliability Engineers focus on designing and improving systems to prevent failures, using data analysis and modeling. Maintenance Engineers perform hands-on repairs and upkeep of equipment to ensure operational continuity. While both roles aim to optimize equipment performance, Reliability Engineers work proactively on system reliability, whereas Maintenance Engineers handle reactive and scheduled maintenance tasks.
What does a reliability engineer do?
As a reliability engineer, your duties are to test and evaluate the manufacturing of products and components and ensure that the procedures are efficient and do not lead to abnormally high maintenance or operational costs. Your other responsibilities are to find solutions to product reliability risks. You may manage risk in a supply chain, develop loss prevention strategies, and track the entire lifecycle of product development, from building prototypes to moving a product into full-scale production. You analyze information from department heads and recommend strategies to reduce risk and ensure that the product works reliably.
What is a reliability engineer?
- Internship Reliability Engineer
- Site Reliability Engineer Intern
- Full Time Site Reliability Engineer Remote
- Senior Site Reliability Engineer
- Entry Level Site Reliability Engineer
- Remote Reliability Engineer
- Senior Reliability Engineer
- Part Time Site Reliability Engineer
- Remote Site Reliability Engineer Intern

Full-time
Retirement
Posted 16 days ago
Job description
Choose a workplace that empowers your impact.
Join a global workplace where employees thrive. One that embraces diversity of thought, expertise and experience. A place where you can personalize your employee journey to be - and deliver - your best.
We are a purpose-driven, dynamic and sustainable pension plan. An industry leading global investor with teams in Toronto to London, New York, Singapore, Sydney and other major cities across North America and Europe. We embody the values of our 665,000 members, placing their best interests at the heart of everything we do.
Join us to accelerate your growth & development, prioritize wellness, build connections, and support the communities where we live and work.
Don't just work anywhere - come build tomorrow together with us.
Know someone at OMERS or Oxford Properties? Great! If you're referred, have them submit your name through Workday first. Then, watch for a unique link in your email to apply.
Role Summary
The Lead, Site Reliability Engineering ensures monitoring and analysis is conducted to guarantee the ongoing stability of all systems. The Lead is also expected to implement and maintain the infrastructure and tools that are necessary to manage the software development process.
Reporting to the SRE, Developer Platform Engineering, the Lead, DEV Platform Support Engineer will play a critical role in supporting and enhancing our Azure-based developer platforms and internal applications. This position combines platform engineering, site reliability engineering, and production support responsibilities to ensure applications remain reliable, secure, and easy to operate.
This is an excellent opportunity for someone who is passionate about Site Reliability Engineering and enjoys improving the reliability, availability, performance, and operability of cloud-native platforms. The successful candidate will apply SRE practices such as incident response, observability, automation, runbook development, root-cause analysis, and continuous reliability improvement while partnering with cross-functional teams to onboard, deploy, monitor, and support applications across the organization.
You Will Be Responsible For
Monitor, troubleshoot, and support applications and developer platform services across DEV, UAT, and PROD environments.
Respond to incidents, lead triage activities, and work closely with SRE, Platform, Network, Security, and application teams to restore service and resolve issues.
Support deployments, release activities, change management, and CI/CD pipelines, including GitHub Actions workflows.
Check, troubleshoot, and resolve user and developer access issues, including Azure AD groups, SSO, application permissions, and firewall rules.
Configure and support platform components such as Azure Container Apps, App Registrations, Key Vault, DNS, certificates, networking, and shared cloud services.
Investigate performance, reliability, and availability issues using Datadog, Azure Monitor, Log Analytics, and related observability tools.
Support onboarding of new applications and teams to the DEV platform by helping with setup, access, deployment readiness, monitoring, and operational handover.
Develop and maintain runbooks, support procedures, knowledge articles, and operational documentation to improve support effectiveness and knowledge sharing.
Contribute to automation and continuous improvement initiatives that reduce manual effort, improve reliability, and strengthen operational processes.
Provide technical guidance to team members and stakeholders while promoting Site Reliability Engineering and platform support best practices.
Required Skills & Experience
5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, DevOps, or Production Support.
Strong hands-on experience with Microsoft Azure services, including Azure Container Apps, Azure Active Directory (Entra ID), Key Vault, Storage Accounts, Azure SQL, API Management (APIM), and Azure Functions.
Experience supporting production environments, including incident response, troubleshooting, problem management, and operational support processes.
Experience with container technologies and cloud-native application architectures.
Hands-on experience with CI/CD pipelines and deployment automation using GitHub Actions or similar platforms.
Strong understanding of identity, networking, and access management concepts, including SSO, OAuth, application registrations, and security groups.
Experience with observability and monitoring platforms such as Datadog, Azure Monitor, and Log Analytics.
Understanding of cloud networking concepts, including DNS, certificates, firewalls, private endpoints, and network security controls.
Experience with scripting and automation using technologies such as PowerShell, Bash, Azure CLI, Python, or similar tools.
Strong knowledge of operating systems and cloud infrastructure concepts.
Proven ability to work effectively in cross-functional environments and collaborate with technical and business stakeholders.
Strong communication, problem-solving, and organizational skills.
Preferred Skills & Experience
Experience with container apps and Kubernetes container orchestration platforms.
Experience with different pipelines.
Experience supporting enterprise developer platforms or internal platform engineering teams.
Knowledge of Site Reliability Engineering principles, including SLOs, SLIs, error budgets, and reliability engineering practices.
Experience with enterprise API integrations and platform services.
Exposure to AI, LLM, or agent-based technology platforms.
Experience with Azure networking and security best practices in enterprise environments.
Azure, Network, DevOps, or cloud-related certifications.
Post-secondary education in Computer Science, Software Engineering, Information Technology, or a related discipline.
We believe that time together in the office is important for OMERS and Oxford, the strength of our employees, and the work we do for our pension members. In delivering on our pension promise, keeping us connected to our work and each other,our flexible hybrid work guideline requires teams to come in to the office 4 days per week.
This posting is for an existing vacancy.The expected salary range for this position is $86,000.00 - $130,000.00 per year.You may also be eligible to receive an annual Incentive Award pursuant to our Short-term Incentive plan and our Long-Term Incentive plan (if applicable), and to participate in our group benefits and retirement plans - details on these elements of compensation are included within OMERS & Oxford offer letters.
As one of Canada's largest defined benefit pension plans, our people-first culture is at its best when our workforce reflects the communities where we live and work - and the members we proudly serve.
From hire to retire, we are an equal opportunity employer committed to an inclusive, barrier-free recruitment and selection process that extends all the way through your employee experience. This sense of belonging and connection is cultivated up, down and across our global organization thanks to our vast network of Employee Resource Groups with executive leader sponsorship, our Purpose@Work committee and employee recognition programs.
Artificial intelligence (AI) tools are used to support certain stages of the OMERS recruitment process. While AI assists us in our process, human judgment and decision-making remain central to our candidate experience.