The AI SRE team is a focused group of SRE engineers dedicated to making PointClickCare's AI and ML platforms reliable, secure, and operationally excellent - from data processing and ML workspaces to ...
The AI SRE team is a focused group of SRE engineers dedicated to making PointClickCare's AI and ML platforms reliable, secure, and operationally excellent - from data processing and ML workspaces to ...
Reporting to the SRE, Developer Platform Engineering , the Lead, DEV Platform Support Engineer will play a critical role in supporting and enhancing our Azure-based developer platforms and internal ...
Reporting to the SRE, Developer Platform Engineering , the Lead, DEV Platform Support Engineer will play a critical role in supporting and enhancing our Azure-based developer platforms and internal ...
Corporate Reliability Engineering Co-Op
CA$22 - CA$25/hr
Supporting Reliability Engineer Data collection and standardization of Failure modes and the Preventive (PM) and Predictive (PdM) Maintenance tasks to maintain the assets * Accurately create ...
Corporate Reliability Engineering Co-Op
CA$22 - CA$25/hr
Supporting Reliability Engineer Data collection and standardization of Failure modes and the Preventive (PM) and Predictive (PdM) Maintenance tasks to maintain the assets * Accurately create ...
Reliability Engineering Specialist
Brantford, ON ยท On-site
CA$95K - CA$125K/yr
As a Reliability Engineering Specialist, you will play a key role in improving asset reliability, reducing downtime, and optimizing maintenance strategies across the site. Working closely with ...
Reliability Engineering Specialist
Brantford, ON ยท On-site
CA$95K - CA$125K/yr
As a Reliability Engineering Specialist, you will play a key role in improving asset reliability, reducing downtime, and optimizing maintenance strategies across the site. Working closely with ...
Head of Platform Engineering, Reliability & Control
Toronto, ON ยท Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
Head of Platform Engineering, Reliability & Control
Toronto, ON ยท Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
Site Reliability Engineer
Vaughan, ON ยท On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Site Reliability Engineer
Vaughan, ON ยท On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
As an Application Support Engineer for our Quantitative Investment Services, you'll take ownership of the reliability and performance of mission-critical trading platforms. This is more than a ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
As an Application Support Engineer for our Quantitative Investment Services, you'll take ownership of the reliability and performance of mission-critical trading platforms. This is more than a ...
Site Reliability Engineer
Toronto, ON ยท On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Site Reliability Engineer
Toronto, ON ยท On-site
The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON ยท Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service during incidents, and minimise customer impact. When issues occur, the role focuses not only on ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON ยท Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service during incidents, and minimise customer impact. When issues occur, the role focuses not only on ...
We follow the Network Reliability Engineer (NRE) model: continuous observation, rapid capacity adaptation, and relentless automation so that network capabilities are delivered as APIs, not tickets.
We follow the Network Reliability Engineer (NRE) model: continuous observation, rapid capacity adaptation, and relentless automation so that network capabilities are delivered as APIs, not tickets.
Director, Reliability Engineering
Toronto, ON ยท On-site
Job Summary The Director, Reliability Engineering is responsible for implementing and leading strategies to maximize uptime, performance and lifespan of assets, and reducing likelihood of failure by ...
Director, Reliability Engineering
Toronto, ON ยท On-site
Job Summary The Director, Reliability Engineering is responsible for implementing and leading strategies to maximize uptime, performance and lifespan of assets, and reducing likelihood of failure by ...
Senior Site Reliability Engineer
Waterloo, ON ยท On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Senior Site Reliability Engineer
Waterloo, ON ยท On-site
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON ยท Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
Quick apply
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON ยท Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
We're searching for a Senior Site Reliability Engineer who not only brings deep technical expertise but also thrives on engaging with customers, partners, and cross-functional teams. In this customer ...
Senior Site Reliability Engineer
Toronto, ON ยท On-site
We're searching for a Senior Site Reliability Engineer who not only brings deep technical expertise but also thrives on engaging with customers, partners, and cross-functional teams. In this customer ...
Elevate cloud platform reliability with Manulife as a Senior Platform Reliability Engineer. Lead Azure ecosystem enhancements, focusing on Kubernetes, CI/CD pipelines, and AI technologies. In this ...
Elevate cloud platform reliability with Manulife as a Senior Platform Reliability Engineer. Lead Azure ecosystem enhancements, focusing on Kubernetes, CI/CD pipelines, and AI technologies. In this ...
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Quick apply
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, production readiness standards, and operational acceptance criteria for ...
Senior Platform Reliability Engineer
Toronto, ON ยท On-site
CA$113K - CA$163K/yr
Senior Platform Reliability Engineer Join Manulife Global Wealth & Asset Management (GWAM) and help power critical cloud platforms that enable data, analytics, and AI across the organization. We're ...
Senior Platform Reliability Engineer
Toronto, ON ยท On-site
CA$113K - CA$163K/yr
Senior Platform Reliability Engineer Join Manulife Global Wealth & Asset Management (GWAM) and help power critical cloud platforms that enable data, analytics, and AI across the organization. We're ...
Reliability Engineer information
What are some typical challenges reliability engineers face when implementing preventive maintenance strategies?
How much do reliability engineers get paid?
What are the key skills and qualifications needed to thrive as a reliability engineer, and why are they important?
What is the difference between Reliability Engineer vs Maintenance Engineer?
| Aspect | Reliability Engineer | Maintenance Engineer |
|---|---|---|
| Credentials | Typically requires engineering degree, certifications in reliability or asset management | Often requires engineering or technical diploma, certifications in maintenance or equipment repair |
| Work Environment | Focuses on analysis, design, and improvement of systems for reliability | Hands-on maintenance, repair, and troubleshooting of equipment |
| Industry Usage | Common in manufacturing, energy, aerospace, and industrial sectors | Prevalent in manufacturing, facilities, and industrial plants |
Reliability Engineers focus on designing and improving systems to prevent failures, using data analysis and modeling. Maintenance Engineers perform hands-on repairs and upkeep of equipment to ensure operational continuity. While both roles aim to optimize equipment performance, Reliability Engineers work proactively on system reliability, whereas Maintenance Engineers handle reactive and scheduled maintenance tasks.
What does a reliability engineer do?
As a reliability engineer, your duties are to test and evaluate the manufacturing of products and components and ensure that the procedures are efficient and do not lead to abnormally high maintenance or operational costs. Your other responsibilities are to find solutions to product reliability risks. You may manage risk in a supply chain, develop loss prevention strategies, and track the entire lifecycle of product development, from building prototypes to moving a product into full-scale production. You analyze information from department heads and recommend strategies to reduce risk and ensure that the product works reliably.
What is a reliability engineer?
- Internship Reliability Engineer
- Site Reliability Engineer Intern
- Full Time Site Reliability Engineer Remote
- Senior Site Reliability Engineer
- Entry Level Site Reliability Engineer
- Remote Reliability Engineer
- Senior Reliability Engineer
- Part Time Site Reliability Engineer
- Remote Site Reliability Engineer Intern

Full-time
Medical, Life, Retirement, PTO
Re-posted 17 days ago
Job description
The AI SRE team is a focused group of SRE engineers dedicated to making PointClickCare's AI and ML platforms reliable, secure, and operationally excellent - from data processing and ML workspaces to model serving and labeling systems. We treat reliability as a product - prioritizing observability, automation, and safe operations so that data scientists and ML engineers can focus on building AI capabilities that improve patient care. You will spend a significant portion of your time hands-on - building automation, designing guardrails, hardening platforms, and leading incident response across cloud environments like Databricks, Azure AI suites. The team collaborates closely with research, platform, data, and security teams across the AI engineering organization.
Job Summary:
AI SRE exists to ensure PointClickCare's AI platforms run safely, reliably, and efficiently - protecting patient data while enabling teams to move fast with confidence. We solve complex cross-cutting reliability and security problems through well-designed automation, SLOs, and operational guardrails - so that product and research teams can focus on delivering AI-driven value to clinicians and patients. We own the infrastructure operability of AI data processing, ML workspaces, labeling systems, and model serving - and we drive the infrastructure observability, incident response, compliance controls, and cost optimization that keep those platforms healthy through sound SRE practices and a security-first mindset
Key responsibilities:
- Own service level objectives, error budgets, and reliability targets for the infrastructure underpinning cloud-based platforms - ensuring infrastructure observability (metrics, logs, traces), alert quality, and telemetry completeness across platform components and serving endpoints
- Design, build, and maintain infrastructure-as-code, operational automation, and change control workflows for AI/ML platforms - with a focus on repeatability, consistency, and toil reduction
- Implement and maintain platform security controls - including network segmentation, secrets management, encryption, and data protection safeguards - aligned to compliance requirements and partnering with security teams to respond to emerging risks
- Lead incident response and blameless postmortems; validate backup/restore and disaster recovery processes; conduct game days and resiliency testing to harden platform and infrastructure reliability
- Mentor engineers, influence design reviews, and collaborate across engineering teams to improve platform resiliency, cost efficiency, capacity planning, and operational standards
- 5+ years in SRE, platform engineering, or infrastructure roles supporting production cloud environments and mission-critical applications
- Strong proficiency with observability - metrics, logging, distributed tracing, SLI/SLO frameworks - and production ownership including incident response, blameless postmortems, and on-call operations
- Strong proficiency with Infrastructure as Code (Terraform), GitOps practices, and CI/CD for infrastructure and platform changes
- Working proficiency with cloud platform administration - compute, networking, storage, and operating managed data or AI/ML platform services in production (e.g., Databricks, Azure ML, or Kubernetes-hosted infrastructure)
- Working proficiency with platform security - network segmentation, secrets management, encryption at rest and in transit, and key management
- Strong programming skills for automation, operational tooling, and infrastructure management
- Strong communication and documentation skills - able to write runbooks, lead postmortems, influence operational standards across teams, and translate technical complexity for diverse audiences
- Experience with disaster recovery planning, multi-region patterns, and capacity or cost optimization (FinOps)
- ย Working knowledge of container orchestration (Kubernetes), progressive delivery patterns (blue/green, canary), and data lineage tooling
- Working knowledge of container orchestration (Kubernetes), progressive delivery patterns (blue/green, canary), and data lineage tooling
- Experience in healthcare, life sciences, or other highly regulated industries with data privacy requirements