At its heart, the Smile platform enables people and organizations to better manage healthcare data ... The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability ...
At its heart, the Smile platform enables people and organizations to better manage healthcare data ... The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability ...
Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is ... Participate in the on-call rotation, and incident management Required Knowledge, Skills, and ...
Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team is ... Participate in the on-call rotation, and incident management Required Knowledge, Skills, and ...
The SRE will be responsible for aligning client requirements to standard features as well as ... Manage and maintain pipelines for client onboarding and live clients on the Azure platform. * Work ...
The SRE will be responsible for aligning client requirements to standard features as well as ... Manage and maintain pipelines for client onboarding and live clients on the Azure platform. * Work ...
Site Reliability Engineer - Interface & Connectivity
Toronto, ON ยท Hybrid
CA$68K - CA$103K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Site Reliability Engineer, you will join one of our Interface ... Support certificate lifecycle management and ensure secure communication protocols are maintained
Site Reliability Engineer - Interface & Connectivity
Toronto, ON ยท Hybrid
CA$68K - CA$103K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Site Reliability Engineer, you will join one of our Interface ... Support certificate lifecycle management and ensure secure communication protocols are maintained
Technical Expertise- Manage and optimize automated application deployments across multiple ... ) methodologies and technologies. Participate in defining standard reliability and resilience ...
New
Technical Expertise- Manage and optimize automated application deployments across multiple ... ) methodologies and technologies. Participate in defining standard reliability and resilience ...
New
Join RBC as the ATM Lead Site Reliability Engineer and lead the reliability, scalability, and ... Manage deployment and pre-production systems ensuring compliance with established standards
New
Join RBC as the ATM Lead Site Reliability Engineer and lead the reliability, scalability, and ... Manage deployment and pre-production systems ensuring compliance with established standards
New
Technical Expertise- Manage and optimize automated application deployments across multiple ... ) methodologies and technologies. Participate in defining standard reliability and resilience ...
New
Technical Expertise- Manage and optimize automated application deployments across multiple ... ) methodologies and technologies. Participate in defining standard reliability and resilience ...
New
Sr. Site Reliability Engineering - Azure Platform
Toronto, ON ยท Hybrid
CA$99K - CA$149K/yr
... E best practices. WHAT YOU WILL BE RESPONSIBLE FOR * Infrastructure Management: Provisioning, configuring, and managing Azure resources (Virtual Machines, Storage, Virtual Networks). * Deploy ...
Sr. Site Reliability Engineering - Azure Platform
Toronto, ON ยท Hybrid
CA$99K - CA$149K/yr
... E best practices. WHAT YOU WILL BE RESPONSIBLE FOR * Infrastructure Management: Provisioning, configuring, and managing Azure resources (Virtual Machines, Storage, Virtual Networks). * Deploy ...
Expert SRE
Mississauga, ON ยท Hybrid
We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ... Manage and validate application deployments across staging and production environments using ...
Expert SRE
Mississauga, ON ยท Hybrid
We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ... Manage and validate application deployments across staging and production environments using ...
Prod Support - SRE
Toronto, ON ยท On-site
Manage and maintain the underlying infrastructure, including servers, networks, and cloud resources ... Work closely with developers, operations teams, and other stakeholders to ensure system reliability ...
Prod Support - SRE
Toronto, ON ยท On-site
Manage and maintain the underlying infrastructure, including servers, networks, and cloud resources ... Work closely with developers, operations teams, and other stakeholders to ensure system reliability ...
The Lead Platform Reliability Engineer (PRE) ensures the stability, performance, and scalability of ... Manage clusters, networks, storage, and policies via Terraform/Ansible; prevent configuration drift.
New
The Lead Platform Reliability Engineer (PRE) ensures the stability, performance, and scalability of ... Manage clusters, networks, storage, and policies via Terraform/Ansible; prevent configuration drift.
New
... reliability ... Develop and manage a national spare parts program, including CAPEX and inventory * You support ...
... reliability ... Develop and manage a national spare parts program, including CAPEX and inventory * You support ...
As Engineering Manager of the Customer Care Platform, you will own the technical strategy for this ... E practices, and the Casa UI component library and design system (CUI), a high-leverage, cross ...
As Engineering Manager of the Customer Care Platform, you will own the technical strategy for this ... E practices, and the Casa UI component library and design system (CUI), a high-leverage, cross ...
... reliability ... Develop and manage a national spare parts program, including CAPEX and inventory * You support ...
... reliability ... Develop and manage a national spare parts program, including CAPEX and inventory * You support ...
Role Overview: We're looking for a Site Reliability Engineer to improve the reliability, resilience ... Managing incidents using your technical know-how to involve the appropriate teams and automate away ...
Role Overview: We're looking for a Site Reliability Engineer to improve the reliability, resilience ... Managing incidents using your technical know-how to involve the appropriate teams and automate away ...
The Team The Auth0 Platform Tools team owns the incident management tooling, Slack-based tooling, StatusPage, and local development environments that Auth0 engineers rely on every day. That includes ...
New
The Team The Auth0 Platform Tools team owns the incident management tooling, Slack-based tooling, StatusPage, and local development environments that Auth0 engineers rely on every day. That includes ...
New
Sr. Observability Engineering
Mississauga, ON ยท On-site +1
Strong working knowledge of SRE principles: SLOs, SLAs, error budgets, alerting philosophy, and incident management. * Scripting/automation proficiency in Python, Go, Bash, or equivalent, with ...
Sr. Observability Engineering
Mississauga, ON ยท On-site +1
Strong working knowledge of SRE principles: SLOs, SLAs, error budgets, alerting philosophy, and incident management. * Scripting/automation proficiency in Python, Go, Bash, or equivalent, with ...
Sr. Observability Engineering
Mississauga, ON ยท On-site +1
Strong working knowledge of SRE principles: SLOs, SLAs, error budgets, alerting philosophy, and incident management. * Scripting/automation proficiency in Python, Go, Bash, or equivalent, with ...
Sr. Observability Engineering
Mississauga, ON ยท On-site +1
Strong working knowledge of SRE principles: SLOs, SLAs, error budgets, alerting philosophy, and incident management. * Scripting/automation proficiency in Python, Go, Bash, or equivalent, with ...
You will influence how teams design systems, validate changes, observe production behavior, manage ... Define and drive BuildOps' technical strategy for engineering quality, production reliability, and ...
You will influence how teams design systems, validate changes, observe production behavior, manage ... Define and drive BuildOps' technical strategy for engineering quality, production reliability, and ...
Senior DevOps Engineer
Toronto, ON ยท On-site
... Reliability Engineering, or Cloud Infrastructure role. * Strong experience with AWS and GCP data services, including Kinesis, Glue, Pub/Sub, and Dataflow. * Proficiency in deploying and managing ...
Senior DevOps Engineer
Toronto, ON ยท On-site
... Reliability Engineering, or Cloud Infrastructure role. * Strong experience with AWS and GCP data services, including Kinesis, Glue, Pub/Sub, and Dataflow. * Proficiency in deploying and managing ...
Reliability Engineer Manager information
What does a reliability engineer manager do?
What are some common challenges reliability engineer managers face when balancing long-term reliability improvements with immediate operational demands?
What are the key skills and qualifications needed to thrive as a reliability engineer manager?
What is the difference between Reliability Engineer Manager vs Reliability Engineer?
| Aspect | Reliability Engineer | Reliability Engineer Manager |
|---|---|---|
| Required Credentials | Bachelor's in Engineering or related field; certifications like CRC, CRE | Same as Reliability Engineer, plus leadership experience |
| Work Environment | Design, analyze, and improve system reliability; often in teams | Oversees Reliability Engineers; manages projects and teams |
| Employer & Industry Usage | Manufacturing, aerospace, energy, automotive | Same industries, with added managerial responsibilities |
| Common Search & Comparison | Focuses on technical skills and hands-on reliability tasks | Focuses on leadership, team management, and strategic planning |
The main difference between a Reliability Engineer and a Reliability Engineer Manager lies in their responsibilities. The Reliability Engineer focuses on technical analysis and system improvements, while the Reliability Engineer Manager oversees teams, manages projects, and develops strategies to enhance reliability across the organization.

Cloud Performance Engineering - Site Reliability Engineer ( Remote Canada)
Toronto, ON โข Remote
Full-time
Medical, Life, Retirement, PTO
Posted 29 days ago
Job description
The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health, its clients, and partners.
This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.
- Collaborate with our Security Operations teams to define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
- Develop, implement, and coordinate a multi-tenant approach around service offerings for databases, container platforms, authentication, certificates, and product registries.
- Design, develop, and maintain cloud performance testing strategies, frameworks, and environments to validate application scalability, reliability, and resiliency.
- Develop and automate load, stress, spike, and endurance (soak) testing as part of CI/CD pipelines.
- Analyze application and infrastructure performance to identify bottlenecks and recommend performance optimizations across cloud-native services.
- Develop and maintain cost and utilization tracking and attribution processes across Cloud Service Providers.
- Create documentation detailing Cloud Service Provider offerings, implementation patterns, and best practices.
- Develop and maintain technical relationships with our core Cloud Service Providers.
- Implement and maintain secure, scalable infrastructure platforms for delivering cloud services.
- Ensure internal and external SLAs are consistently met or exceeded, while continuously monitoring and improving system performance, reliability, and availability.
- Create tools for automating deployment, monitoring, and platform operations.
- Implement and manage observability solutions (logging, metrics, tracing) using OpenTelemetry, Prometheus, Grafana, Azure Monitor, and related technologies to provide actionable performance insights.
- Plan and execute chaos engineering experiments to evaluate and improve application resiliency and fault tolerance.
- 5+ years of experience with Cloud Service Providers and best practices around implementation and configuration, preferably managing Azure environments supporting SaaS products.
- Experience working across multiple cloud providers (Azure required; AWS and/or Google Cloud Platform considered an asset).
- Strong experience in Cloud Performance Engineering, including performance analysis, capacity planning, scalability testing, and optimization of distributed cloud-native applications.
- Proven experience working with microservices architecture, with a strong focus on Java-based services.
- Experience applying Chaos Engineering practices to evaluate and improve system resiliency.
- Strong experience designing and executing performance testing strategies, including load, stress, spike, and endurance (soak) testing, to validate application scalability and defined latency and error-rate thresholds.
- Hands-on experience with performance testing tools such as JMeter, Gatling, Azure Load Testing, or k6.
- Experience validating application services sustaining 500+ transactions per second (TPS) while meeting defined performance objectives.
- Hands-on experience deploying and managing containerized applications using Docker and Kubernetes, including autoscaling and performance optimization.
- Experience using Terraform to provision and manage cloud infrastructure using Infrastructure as Code (IaC).
- Experience tuning Kafka (partitioning, consumer group sizing, throughput/latency trade-offs) and other messaging/queueing platforms to sustain target transaction rates.
- Hands-on experience implementing and using observability platforms including OpenTelemetry, Prometheus, Grafana, Azure Monitor, Application Insights, and Log Analytics.
- Proven experience with Security and Compliance (SOC 2, HIPAA, ISO 27001) best practices and implementing controls that support high-velocity software delivery teams.