1

Reliability Engineer Manager Jobs in Toronto, ON

Lead major incident management and stakeholder communication. * Mentor junior engineers and promote SRE best practices. * Ensure system reliability and uptime by designing, implementing, and ...

... and access management, cloud computing services/integrations, and data analytics technologies. Responsibilities of the Cloud Service Reliability Engineer: * Establishes technology product ...

SRE is part of a global organization that leverages the latest technology to communicate with our ... You'll be a key voice in observability, change management, and service scalability, providing ...

SRE Evaluator - Incident Management

Toronto, ON ยท Remote

CA$80 - CA$120/hr

Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...

Site Reliability Engineer

Toronto, ON ยท On-site +1

CA$125K - CA$250K/yr

Improve provisioning, configuration management, testing, and deployment automation * Help plan ... site reliability engineering, infrastructure engineering, systems engineering, or a related ...

Site Reliability Engineer (.Net)

Toronto, ON ยท Hybrid

CA$100K - CA$125K/yr

As a Site Reliability Engineer, you will play a crucial role in enhancing the reliability ... supplier management, tax compliance, and treasury. Tipalti partners with leading financial ...

next page

Showing results 1-20

Reliability Engineer Manager information

What does a reliability engineer manager do?

A Reliability Engineer Manager oversees teams responsible for improving the reliability and performance of systems, machinery, or processes within an organization. They develop maintenance strategies, lead root cause analyses of failures, and implement best practices to minimize downtime and costs. Additionally, they collaborate with other departments to ensure that reliability goals align with business objectives and compliance standards. Their role is crucial in industries such as manufacturing, energy, and technology, where system uptime and safety are critical.

What are some common challenges reliability engineer managers face when balancing long-term reliability improvements with immediate operational demands?

Reliability Engineer Managers often need to prioritize urgent maintenance issues while also driving long-term reliability initiatives. Balancing these competing demands can be challenging, as immediate equipment failures may require quick fixes that temporarily interrupt ongoing improvement projects. Effective managers work closely with operations, maintenance, and engineering teams to communicate priorities, allocate resources, and implement sustainable solutions that address root causes rather than just symptoms. This role typically involves using data-driven decision-making and fostering a culture of proactive maintenance and continuous improvement.

What are the key skills and qualifications needed to thrive as a reliability engineer manager?

To thrive as a Reliability Engineer Manager, you need a strong background in engineering principles, reliability analysis, and maintenance strategies, typically supported by a degree in engineering and experience in reliability roles. Familiarity with reliability-centered maintenance (RCM), failure mode and effects analysis (FMEA), and asset management software such as SAP or Maximo is common, along with certifications like Certified Reliability Engineer (CRE). Leadership, problem-solving, and effective communication are vital soft skills for managing teams and driving cross-functional initiatives. These competencies are crucial for minimizing downtime, optimizing equipment performance, and ensuring long-term operational efficiency.

What is the difference between Reliability Engineer Manager vs Reliability Engineer?

AspectReliability EngineerReliability Engineer Manager
Required CredentialsBachelor's in Engineering or related field; certifications like CRC, CRESame as Reliability Engineer, plus leadership experience
Work EnvironmentDesign, analyze, and improve system reliability; often in teamsOversees Reliability Engineers; manages projects and teams
Employer & Industry UsageManufacturing, aerospace, energy, automotiveSame industries, with added managerial responsibilities
Common Search & ComparisonFocuses on technical skills and hands-on reliability tasksFocuses on leadership, team management, and strategic planning

The main difference between a Reliability Engineer and a Reliability Engineer Manager lies in their responsibilities. The Reliability Engineer focuses on technical analysis and system improvements, while the Reliability Engineer Manager oversees teams, manages projects, and develops strategies to enhance reliability across the organization.

What are the most commonly searched types of Reliability Engineer jobs in Toronto, ON? The most popular types of Reliability Engineer jobs in Toronto, ON are:
Infographic showing various Reliability Engineer Manager job openings in Toronto, ON as of July 2026, with employment types broken down into 94% Full Time, 3% Part Time, and 3% Contract. Highlights an 86% Physical, 7% Hybrid, and 7% Remote job distribution.

Senior Site Reliability Engineer (SRE)

Acquird.io

Toronto, ON โ€ข On-site

$130 - $180/hr

Other

Medical

Re-posted 20 days ago


Job description

A Few Notes:

  • Profitable B2B SaaS company, teams are based out of North America

  • Role is 95% remote in Toronto (we meetup 1x a month).

  • Must be able to legally work in Canada (visa or sponsorship won't be provided)

  • Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer

  • Our main Cloud Platform is Azure (those with Azure will be prioritized first)

About Us:

We're one of the top retail analytic platforms that help marketing teams/brands understand their retail data and run targeted media campaigns without writing code. We help our clients better understand their customers and improve their ROI on campaigns. One of our main customers is Home Depot.

  • Modern Cloud Stack: Azure is our primary cloud. CI/CD, Containerization, Distributed computing.

About You:

We are looking for an outstanding Senior SRE/Cloud Engineer who wants to play a key role in running our Cloud Ops (ensuring uptime, reliability, CI/CD, Containerization, and automation)

Example Responsibilities:

  • Collaborate with software engineering teams to design, implement, and maintain CI/CD pipelines, enabling rapid and reliable software releases.

  • Automate and optimize our infrastructure provisioning, configuration, and management processes using industry-standard tools and best practices.

  • Implement and manage containerization and orchestration technologies to enhance scalability and resource utilization.

  • Take ownership of the end-to-end availability and performance of our cloud infrastructure; proactively identifying potential issues, and implementing automation to prevent the recurrence of problems.

  • Participate in an on-call rotation, ensuring our systems remain stable and responsive even during off-hours.

  • Lead the development, implementation, and achievement of service-level objectives that are instrumental in maintaining product reliability.

  • Maintain and enhance version control systems and repositories for codebase management.

  • Steer and drive the SRE / DevOps roadmap, assuming full ownership while actively engaging in negotiation and strategic planning to ensure its successful execution.

  • Stay current with industry trends, emerging technologies, and best practices in SRE, DevOps, and automation.

Qualifications:

  • 5 years plus of experience as a Site Reliability Engineer or DevOps Engineer, working with software and infrastructure.

  • Bachelorโ€™s degree in Computer Science, a related technical field, or equivalent practical experience.

  • Experience in one or more of the following: Python, Javascript, Ruby, Groovy, PHP, or Bash.

  • Experience in one of the cloud platforms: Azure, AWS, or GCP. (Azure is our main Cloud Platform)

  • Led and built cloud infrastructure projects (zero to one, automation, start up experience)

  • Nice to have: Experience with high availability systems.

  • Nice to have: Experience troubleshooting and debugging production code.

  • Nice to have: Experience with application deployment and data pipelines.

  • Nice to have: Understanding of distributed computing systems.

  • Nice to have: Experience with Snowflake and/or relational databases

Target FT Salary Range:

  • $130,000 - $180,000* base CDN a year, with Equity & Health Benefits

  • *Comp range is higher for Staff level

Benefits:

Meritocracy

- Leadership opportunities as we scale

- Equity/Options grants

Generous Time Off

- Flexible remote work policies

Platinum Benefits

- Solid Health Insurance Plan on Day 1

Learn and Grow

- Coaching/mentoring for your professional development

#J-18808-Ljbffr