1

Reliability Engineer Manager Jobs in Toronto, ON

We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ... Manage and validate application deployments across staging and production environments using ...

Manage and maintain the underlying infrastructure, including servers, networks, and cloud resources ... Work closely with developers, operations teams, and other stakeholders to ensure system reliability ...

Strong working knowledge of SRE principles: SLOs, SLAs, error budgets, alerting philosophy, and incident management. * Scripting/automation proficiency in Python, Go, Bash, or equivalent, with ...

Strong working knowledge of SRE principles: SLOs, SLAs, error budgets, alerting philosophy, and incident management. * Scripting/automation proficiency in Python, Go, Bash, or equivalent, with ...

Showing results 41-60

Reliability Engineer Manager information

What does a reliability engineer manager do?

A Reliability Engineer Manager oversees teams responsible for improving the reliability and performance of systems, machinery, or processes within an organization. They develop maintenance strategies, lead root cause analyses of failures, and implement best practices to minimize downtime and costs. Additionally, they collaborate with other departments to ensure that reliability goals align with business objectives and compliance standards. Their role is crucial in industries such as manufacturing, energy, and technology, where system uptime and safety are critical.

What are some common challenges reliability engineer managers face when balancing long-term reliability improvements with immediate operational demands?

Reliability Engineer Managers often need to prioritize urgent maintenance issues while also driving long-term reliability initiatives. Balancing these competing demands can be challenging, as immediate equipment failures may require quick fixes that temporarily interrupt ongoing improvement projects. Effective managers work closely with operations, maintenance, and engineering teams to communicate priorities, allocate resources, and implement sustainable solutions that address root causes rather than just symptoms. This role typically involves using data-driven decision-making and fostering a culture of proactive maintenance and continuous improvement.

What are the key skills and qualifications needed to thrive as a reliability engineer manager?

To thrive as a Reliability Engineer Manager, you need a strong background in engineering principles, reliability analysis, and maintenance strategies, typically supported by a degree in engineering and experience in reliability roles. Familiarity with reliability-centered maintenance (RCM), failure mode and effects analysis (FMEA), and asset management software such as SAP or Maximo is common, along with certifications like Certified Reliability Engineer (CRE). Leadership, problem-solving, and effective communication are vital soft skills for managing teams and driving cross-functional initiatives. These competencies are crucial for minimizing downtime, optimizing equipment performance, and ensuring long-term operational efficiency.

What is the difference between Reliability Engineer Manager vs Reliability Engineer?

AspectReliability EngineerReliability Engineer Manager
Required CredentialsBachelor's in Engineering or related field; certifications like CRC, CRESame as Reliability Engineer, plus leadership experience
Work EnvironmentDesign, analyze, and improve system reliability; often in teamsOversees Reliability Engineers; manages projects and teams
Employer & Industry UsageManufacturing, aerospace, energy, automotiveSame industries, with added managerial responsibilities
Common Search & ComparisonFocuses on technical skills and hands-on reliability tasksFocuses on leadership, team management, and strategic planning

The main difference between a Reliability Engineer and a Reliability Engineer Manager lies in their responsibilities. The Reliability Engineer focuses on technical analysis and system improvements, while the Reliability Engineer Manager oversees teams, manages projects, and develops strategies to enhance reliability across the organization.

What are the most commonly searched types of Reliability Engineer jobs in Toronto, ON? The most popular types of Reliability Engineer jobs in Toronto, ON are:
What are popular job titles related to Reliability Engineer Manager jobs in Toronto, ON? For Reliability Engineer Manager jobs in Toronto, ON, the most frequently searched job titles are:
What job categories do people searching Reliability Engineer Manager jobs in Toronto, ON look for? The top searched job categories for Reliability Engineer Manager jobs in Toronto, ON are:
Infographic showing various Reliability Engineer Manager job openings in Toronto, ON as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% In-person job distribution.

Cloud Performance Engineering - Site Reliability Engineer ( Remote Canada)

Smile Digital Health

Toronto, ON โ€ข Remote

Full-time

Medical, Life, Retirement, PTO

Posted 29 days ago


Job description

Working for a company like Smile Digital Health means supporting our mandate for #BetterGlobalHealth. We strive towards this goal every day, and the results can be seen in the impact of our innovative health data platform and data management solutions, which are used in over 20 countries. We were #19 on Deloitte's Technology Fast 50 Ranking for 2024!ย 
ย 
Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform.
ย 
At its heart, the Smile platform enables people and organizations to better manage healthcare data. We helpย generate andย liberate structured healthcare data to ensure effective delivery across care teams and health systems bringing ย #BetterGlobalHealth to patients everyday!

Apply today and find plenty of reasons to SMILE!

The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health, its clients, and partners.

This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.

Responsibilities:
  • Collaborate with our Security Operations teams to define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
  • Develop, implement, and coordinate a multi-tenant approach around service offerings for databases, container platforms, authentication, certificates, and product registries.
  • Design, develop, and maintain cloud performance testing strategies, frameworks, and environments to validate application scalability, reliability, and resiliency.
  • Develop and automate load, stress, spike, and endurance (soak) testing as part of CI/CD pipelines.
  • Analyze application and infrastructure performance to identify bottlenecks and recommend performance optimizations across cloud-native services.
  • Develop and maintain cost and utilization tracking and attribution processes across Cloud Service Providers.
  • Create documentation detailing Cloud Service Provider offerings, implementation patterns, and best practices.
  • Develop and maintain technical relationships with our core Cloud Service Providers.
  • Implement and maintain secure, scalable infrastructure platforms for delivering cloud services.
  • Ensure internal and external SLAs are consistently met or exceeded, while continuously monitoring and improving system performance, reliability, and availability.
  • Create tools for automating deployment, monitoring, and platform operations.
  • Implement and manage observability solutions (logging, metrics, tracing) using OpenTelemetry, Prometheus, Grafana, Azure Monitor, and related technologies to provide actionable performance insights.
  • Plan and execute chaos engineering experiments to evaluate and improve application resiliency and fault tolerance.
Requirements:
  • 5+ years of experience with Cloud Service Providers and best practices around implementation and configuration, preferably managing Azure environments supporting SaaS products.
  • Experience working across multiple cloud providers (Azure required; AWS and/or Google Cloud Platform considered an asset).
  • Strong experience in Cloud Performance Engineering, including performance analysis, capacity planning, scalability testing, and optimization of distributed cloud-native applications.
  • Proven experience working with microservices architecture, with a strong focus on Java-based services.
  • Experience applying Chaos Engineering practices to evaluate and improve system resiliency.
  • Strong experience designing and executing performance testing strategies, including load, stress, spike, and endurance (soak) testing, to validate application scalability and defined latency and error-rate thresholds.
  • Hands-on experience with performance testing tools such as JMeter, Gatling, Azure Load Testing, or k6.
  • Experience validating application services sustaining 500+ transactions per second (TPS) while meeting defined performance objectives.
  • Hands-on experience deploying and managing containerized applications using Docker and Kubernetes, including autoscaling and performance optimization.
  • Experience using Terraform to provision and manage cloud infrastructure using Infrastructure as Code (IaC).
  • Experience tuning Kafka (partitioning, consumer group sizing, throughput/latency trade-offs) and other messaging/queueing platforms to sustain target transaction rates.
  • Hands-on experience implementing and using observability platforms including OpenTelemetry, Prometheus, Grafana, Azure Monitor, Application Insights, and Log Analytics.
  • Proven experience with Security and Compliance (SOC 2, HIPAA, ISO 27001) best practices and implementing controls that support high-velocity software delivery teams.
$110,000 - $125,000 a year
Smile discloses that artificial intelligence (AI) may be used in portions of the recruitment and selection process, such as resume screening or application assessment. All hiring decisions are ultimately made by qualified human decision-makers, and AI tools are used to support - not replace - fair and equitable hiring practices.ย 
ย 
This position is a replacement role, created to support Smile's continued growth and commitment to operational excellence.

Some of the benefits we offer:
* Remote Work Environment
* Flexible Time Away From Work Policy including PTO, Personal and Sick Days
* Competitive Salary and Health/Medical Benefits
* RRSP/TFSA/401K Employee Contribution
* Life and Disability
* Employee Assistance Program
* FHIR Study Program and Skillsoft Learning
* Super HAPI Fun Club

Smile's core values include respect, inclusion, embracing our differences, and celebrating shared values because our people are the foundation of our success. We are big on creating a sense of belonging and empowering each other to bring our authentic selves to work.ย  We are dedicated to fostering a workplace that values diversity, equity, and inclusion.
ย 
We welcome and encourage candidates of all backgrounds to apply. Candidates are encouraged to inform us if they wish to discuss or require accommodations during interviews or while working at Smile.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
apply for this job