The AI SRE team is a focused group of SRE engineers dedicated to making PointClickCare's AI and ML platforms reliable, secure, and operationally excellent - from data processing and ML workspaces to ...
The AI SRE team is a focused group of SRE engineers dedicated to making PointClickCare's AI and ML platforms reliable, secure, and operationally excellent - from data processing and ML workspaces to ...
... (SRE) methodologies and technologies. Participate in defining standard reliability and resilience requirements for infrastructure and application components. Act as a point of escalation for ...
... (SRE) methodologies and technologies. Participate in defining standard reliability and resilience requirements for infrastructure and application components. Act as a point of escalation for ...
... experienced Site Reliability Engineer (SRE)/DevOps Developer to help build, operate, and ... In this role, you will partner closely with Software Engineering, Infrastructure, and Platform ...
... experienced Site Reliability Engineer (SRE)/DevOps Developer to help build, operate, and ... In this role, you will partner closely with Software Engineering, Infrastructure, and Platform ...
Sr. Site Reliability Engineering - Azure Platform
Toronto, ON · Hybrid
CA$99K - CA$149K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Senior Site Reliability Engineer, you will be embedded within ... Configure and support third-party software (e.g. IDP, SFTP) in client environments. * Understand ...
Sr. Site Reliability Engineering - Azure Platform
Toronto, ON · Hybrid
CA$99K - CA$149K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Senior Site Reliability Engineer, you will be embedded within ... Configure and support third-party software (e.g. IDP, SFTP) in client environments. * Understand ...
The SRE will be responsible for aligning client requirements to standard features as well as deploying those features using automated technologies. Our SREs are pivotal in delivering technology ...
The SRE will be responsible for aligning client requirements to standard features as well as deploying those features using automated technologies. Our SREs are pivotal in delivering technology ...
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
Quick apply
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
Quick apply
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
Quick apply
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
Quick apply
We are looking for an experienced Site Reliability Engineer or Platform Operations Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or ...
... digital investigative software that acquires, analyzes, and shares evidence from computers ... Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ...
... digital investigative software that acquires, analyzes, and shares evidence from computers ... Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ...
... digital investigative software that acquires, analyzes, and shares evidence from computers ... Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ...
... digital investigative software that acquires, analyzes, and shares evidence from computers ... Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ...
... digital investigative software that acquires, analyzes, and shares evidence from computers ... Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ...
... digital investigative software that acquires, analyzes, and shares evidence from computers ... Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within ...
Join RBC as the ATM Lead Site Reliability Engineer and lead the reliability, scalability, and ... Oversee end-to-end reliability of the ATM ecosystem (hardware, software, network) ensuring 99.7% ...
Join RBC as the ATM Lead Site Reliability Engineer and lead the reliability, scalability, and ... Oversee end-to-end reliability of the ATM ecosystem (hardware, software, network) ensuring 99.7% ...
As a Site Reliability Engineer on the Application Edge team at Proton, you will help build ... Experience with software development and automation in one or more programming languages (e.g ...
As a Site Reliability Engineer on the Application Edge team at Proton, you will help build ... Experience with software development and automation in one or more programming languages (e.g ...
Experience in Application Support, Production Support, SRE, or Platform Operations roles. Strong understanding of incident, problem, and change management processes. Experience supporting mission ...
Experience in Application Support, Production Support, SRE, or Platform Operations roles. Strong understanding of incident, problem, and change management processes. Experience supporting mission ...
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON · Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
Quick apply
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON · Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
Expert SRE
Mississauga, ON · Hybrid
Who are we? At Finastra, we're a global leader in financial services software, dedicated to ... We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ...
Expert SRE
Mississauga, ON · Hybrid
Who are we? At Finastra, we're a global leader in financial services software, dedicated to ... We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ...
Extensive experience in an SRE, DevOps, or Platform Engineering role. * Comfortable writing Python ... Technology: you'll get the right hardware and the right software you need to do your best work.
Extensive experience in an SRE, DevOps, or Platform Engineering role. * Comfortable writing Python ... Technology: you'll get the right hardware and the right software you need to do your best work.
DevOps / SRE Engineer (Remote)
Toronto, ON · On-site +1
To do that we are eager to add a highly skilled DevOps / SRE Engineer Engineer to our incredible ... This is a senior role working alongside the current backend and frontend software engineering team.
Quick apply
DevOps / SRE Engineer (Remote)
Toronto, ON · On-site +1
To do that we are eager to add a highly skilled DevOps / SRE Engineer Engineer to our incredible ... This is a senior role working alongside the current backend and frontend software engineering team.
To do that we are eager to add a highly skilled DevOps / SRE Engineer Engineer to our incredible ... This is a senior role working alongside the current backend and frontend software engineering team.
Quick apply
To do that we are eager to add a highly skilled DevOps / SRE Engineer Engineer to our incredible ... This is a senior role working alongside the current backend and frontend software engineering team.
Software Engineer Site Reliability Engineer information
What is the difference between Software Engineer Site Reliability Engineer vs DevOps Engineer?
| Aspect | Software Engineer Site Reliability Engineer | DevOps Engineer |
|---|---|---|
| Credentials | Bachelor's in CS or related, sometimes certifications in cloud or SRE practices | Bachelor's in CS, IT, or related, with certifications in cloud, automation, or CI/CD tools |
| Work Environment | Focus on reliability, scalability, and automation within software development teams | Bridge between development and operations, emphasizing automation, deployment, and infrastructure |
| Employer & Industry Usage | Tech companies, cloud providers, large enterprises | Startups, tech firms, organizations adopting DevOps practices |
While both roles focus on automation and system stability, Software Engineer Site Reliability Engineers primarily ensure system reliability and performance, whereas DevOps Engineers focus on streamlining development and deployment processes. The roles often overlap but differ in their core focus areas and daily responsibilities.
What are popular job titles related to Software Engineer Site Reliability Engineer jobs in Ontario?
For Software Engineer Site Reliability Engineer jobs in Ontario, the most frequently searched job titles are:
What job categories do people searching Software Engineer Site Reliability Engineer jobs in Ontario look for?
The top searched job categories for Software Engineer Site Reliability Engineer jobs in Ontario are:
What cities in Ontario are hiring for Software Engineer Site Reliability Engineer jobs?
Cities in Ontario with the most Software Engineer Site Reliability Engineer job openings:
Full-time
Medical, Life, Retirement, PTO
Re-posted 24 days ago
Job description
The AI SRE team is a focused group of SRE engineers dedicated to making PointClickCare's AI and ML platforms reliable, secure, and operationally excellent - from data processing and ML workspaces to model serving and labeling systems. We treat reliability as a product - prioritizing observability, automation, and safe operations so that data scientists and ML engineers can focus on building AI capabilities that improve patient care. You will spend a significant portion of your time hands-on - building automation, designing guardrails, hardening platforms, and leading incident response across cloud environments like Databricks, Azure AI suites. The team collaborates closely with research, platform, data, and security teams across the AI engineering organization.
Job Summary:
AI SRE exists to ensure PointClickCare's AI platforms run safely, reliably, and efficiently - protecting patient data while enabling teams to move fast with confidence. We solve complex cross-cutting reliability and security problems through well-designed automation, SLOs, and operational guardrails - so that product and research teams can focus on delivering AI-driven value to clinicians and patients. We own the infrastructure operability of AI data processing, ML workspaces, labeling systems, and model serving - and we drive the infrastructure observability, incident response, compliance controls, and cost optimization that keep those platforms healthy through sound SRE practices and a security-first mindset
Key responsibilities:
- Own service level objectives, error budgets, and reliability targets for the infrastructure underpinning cloud-based platforms - ensuring infrastructure observability (metrics, logs, traces), alert quality, and telemetry completeness across platform components and serving endpoints
- Design, build, and maintain infrastructure-as-code, operational automation, and change control workflows for AI/ML platforms - with a focus on repeatability, consistency, and toil reduction
- Implement and maintain platform security controls - including network segmentation, secrets management, encryption, and data protection safeguards - aligned to compliance requirements and partnering with security teams to respond to emerging risks
- Lead incident response and blameless postmortems; validate backup/restore and disaster recovery processes; conduct game days and resiliency testing to harden platform and infrastructure reliability
- Mentor engineers, influence design reviews, and collaborate across engineering teams to improve platform resiliency, cost efficiency, capacity planning, and operational standards
- 5+ years in SRE, platform engineering, or infrastructure roles supporting production cloud environments and mission-critical applications
- Strong proficiency with observability - metrics, logging, distributed tracing, SLI/SLO frameworks - and production ownership including incident response, blameless postmortems, and on-call operations
- Strong proficiency with Infrastructure as Code (Terraform), GitOps practices, and CI/CD for infrastructure and platform changes
- Working proficiency with cloud platform administration - compute, networking, storage, and operating managed data or AI/ML platform services in production (e.g., Databricks, Azure ML, or Kubernetes-hosted infrastructure)
- Working proficiency with platform security - network segmentation, secrets management, encryption at rest and in transit, and key management
- Strong programming skills for automation, operational tooling, and infrastructure management
- Strong communication and documentation skills - able to write runbooks, lead postmortems, influence operational standards across teams, and translate technical complexity for diverse audiences
- Experience with disaster recovery planning, multi-region patterns, and capacity or cost optimization (FinOps)
- Working knowledge of container orchestration (Kubernetes), progressive delivery patterns (blue/green, canary), and data lineage tooling
- Working knowledge of container orchestration (Kubernetes), progressive delivery patterns (blue/green, canary), and data lineage tooling
- Experience in healthcare, life sciences, or other highly regulated industries with data privacy requirements