1

Reliability Engineer Manager Jobs in Ontario (NOW HIRING)

The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...

We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ... Create, manage, and support pipelines that the application support teams will be utilizing to ...

The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...

Showing results 41-60

Reliability Engineer Manager information

What does a reliability engineer manager do?

A Reliability Engineer Manager oversees teams responsible for improving the reliability and performance of systems, machinery, or processes within an organization. They develop maintenance strategies, lead root cause analyses of failures, and implement best practices to minimize downtime and costs. Additionally, they collaborate with other departments to ensure that reliability goals align with business objectives and compliance standards. Their role is crucial in industries such as manufacturing, energy, and technology, where system uptime and safety are critical.

What are some common challenges reliability engineer managers face when balancing long-term reliability improvements with immediate operational demands?

Reliability Engineer Managers often need to prioritize urgent maintenance issues while also driving long-term reliability initiatives. Balancing these competing demands can be challenging, as immediate equipment failures may require quick fixes that temporarily interrupt ongoing improvement projects. Effective managers work closely with operations, maintenance, and engineering teams to communicate priorities, allocate resources, and implement sustainable solutions that address root causes rather than just symptoms. This role typically involves using data-driven decision-making and fostering a culture of proactive maintenance and continuous improvement.

What are the key skills and qualifications needed to thrive as a reliability engineer manager?

To thrive as a Reliability Engineer Manager, you need a strong background in engineering principles, reliability analysis, and maintenance strategies, typically supported by a degree in engineering and experience in reliability roles. Familiarity with reliability-centered maintenance (RCM), failure mode and effects analysis (FMEA), and asset management software such as SAP or Maximo is common, along with certifications like Certified Reliability Engineer (CRE). Leadership, problem-solving, and effective communication are vital soft skills for managing teams and driving cross-functional initiatives. These competencies are crucial for minimizing downtime, optimizing equipment performance, and ensuring long-term operational efficiency.

What is the difference between Reliability Engineer Manager vs Reliability Engineer?

AspectReliability EngineerReliability Engineer Manager
Required CredentialsBachelor's in Engineering or related field; certifications like CRC, CRESame as Reliability Engineer, plus leadership experience
Work EnvironmentDesign, analyze, and improve system reliability; often in teamsOversees Reliability Engineers; manages projects and teams
Employer & Industry UsageManufacturing, aerospace, energy, automotiveSame industries, with added managerial responsibilities
Common Search & ComparisonFocuses on technical skills and hands-on reliability tasksFocuses on leadership, team management, and strategic planning

The main difference between a Reliability Engineer and a Reliability Engineer Manager lies in their responsibilities. The Reliability Engineer focuses on technical analysis and system improvements, while the Reliability Engineer Manager oversees teams, manages projects, and develops strategies to enhance reliability across the organization.

What are the most commonly searched types of Reliability Engineer jobs in Ontario?

The most popular types of Reliability Engineer jobs in Ontario are:

What are popular job titles related to Reliability Engineer Manager jobs in Ontario?

For Reliability Engineer Manager jobs in Ontario, the most frequently searched job titles are:

What job categories do people searching Reliability Engineer Manager jobs in Ontario look for?

The top searched job categories for Reliability Engineer Manager jobs in Ontario are:

What cities in Ontario are hiring for Reliability Engineer Manager jobs?

Cities in Ontario with the most Reliability Engineer Manager job openings:

Infographic showing various Reliability Engineer Manager job openings in Ontario as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% In-person job distribution.

Senior Site Reliability Engineer

Magnet Forensics

Toronto, ON

Full-time

Medical, Retirement

Re-posted 15 days ago


Job description

Who We Are; What We Do; Where We're Going
 
Magnet Forensics is a global leader in the development of digital investigative software that acquires, analyzes, and shares evidence from computers, smartphones, tablets, and IoT-related devices. We are continually innovating so our customers can deploy advanced and effective tools to protect their companies, communities, and countries.
 
Serving thousands of customers globally, our solutions are playing a crucial role in modernizing digital investigations, helping investigators fight crime, protect assets, and guard national security.
 
With employees based around the world, Magnet Forensics has been expanding our global presence. As a part of Magnet Forensics, you can expect to make a difference in the world, no matter what role you play. You'll be supported through learning and development, not to mention an incredible team with unbelievable talent and integrity.
 
If you think you would be the right person to join our team working towards this goal, we would love to hear from you! 

Role Overview 
We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly available SaaS platform, a production Kubernetes environment serving law enforcement and government customers globally. 
 
This role requires deep AWS expertise, infrastructure-as-code discipline, and CI/CD best practices. You'll work closely with Application, Platform, and Security teams to drive secure-by-design architectures and improve automation and reliability across our cloud environments. You'll ship infrastructure as code, respond to production incidents with discipline, and drive platform modernization through deliberate roadmap execution.
 
As part of the SaaS-Ops team, you'll work in a high-performing environment where members take ownership of outcomes and operate with a strong sense of trust and autonomy. You'll identify challenges, contribute to solutions, raise concerns proactively, support improvements, and navigate situations requiring timely decision-making. If you're looking for your next challenge where infrastructure quality directly impacts realworld outcomes, this role could be a great fit!
 
 
Note: This role includes participation in an on-call rotation.
What You'll Do
  • Own and operate production Kubernetes clusters (Amazon EKS) including upgrades, scaling, security hardening, and cluster lifecycle management;
  • Design, implement, and maintain infrastructure-as-code using Terraform; contribute to shared module libraries and enforce IaC standards across the team;

  • Manage and evolve Helm chart definitions and ArgoCD GitOps workflows for multi-region SaaS deployments;

  • Operate and maintain observability infrastructure including Grafana, alerts, dashboards, and log pipelines. Act to eliminate noise and surface signal;

  • Contribute to pipeline reliability: identify flaky stages, reduce build times, improve developer experience across CI/CD pipelines;

  • Remediate security vulnerabilities (CVEs) in container images and infrastructure components; participate in compliance work including FedRAMP support activities;

  • Develop and maintain runbooks, change management procedures, and operational documentation;

  • Ensure alignment with internal policies and frameworks such as ISO 27001, SOC2, and NIST;

  • Contribute to AI-assisted tooling and automation (e.g., Claude-based Terraform agents, automated triage tools) as part of the team's operational efficiency roadmap;

  • Participate in on-call incident response rotation; lead or support incident command during active production incidents including root cause analysis and post-incident review.

What We're Looking For
  • 5+ years of industry experience with a trajectory that demonstrates growing depth in cloud infrastructure and SRE practices;

  • Managed production Kubernetes environments at scale: not just deployed workloads, but owned cluster health, upgrades, and failure modes;

  • Responded to production incidents in high-stakes environments where downtime has real consequences;

  • Written and maintained Terraform at the module level, not just as a consumer: understands state, dependencies, and the operational burden of drift;

  • Operated in an environment that uses GitOps: has a good understanding of Helm chart organization, ArgoCD app-of-apps patterns, or equivalent;

  • Balanced reactive operational work with proactive roadmap delivery; knows how to protect time for improvements while keeping production stable;

  • Worked with observability as a first-class discipline: built meaningful dashboards, eliminated alert fatigue, and used metrics to make operational decisions;

  • Contributed to security hardening in a regulated or compliance-adjacent environment: FedRAMP, SOC 2, or similar frameworks are a strong asset.

Compensation & Benefits 
The Compensation range is for the primary location for which the job is posted. Please note that the actual compensation may vary depending on location and job-related factors such as qualifications, experience, knowledge and skills. If you are applying for this role outside of the primary location and you are selected for an interview, the Talent Acquisition Partner can share more information with you.  If the compensation structure for the role includes an incentive component (i.e. most Sales roles) the range below represents total target compensation (TTC) (base salary + variable).
 
$110,000 - $160,000 (CDN) a year 
 
Position Type: Current Vacancy 

Magnet is proud to offer benefits such as: 
 
- Generous time off policies 
- Competitive compensation 
- Volunteer opportunities  
- Reward and recognition programs   
- Employee committees & resource groups  
- Healthcare and retirement benefits 
 
Indicators of Success
 
We're looking for someone who checks off most, but not all, of the boxes listed in "skills and experiences".  It's more important to us to find candidates who can display indicators of success through skills they have developed and experiences they have been a part of, than to find folks who have "been there, done that".  We want to be part of your development journey, and we'll learn as much from you as you learn from us. 
 
How We Work
 
At Magnet Forensics, we take a hybrid-flexible approach to support your productivity and work-life balance. If you're within a comfortable travel distance to one of our offices, you'll occasionally join us in person. How often you'll come in depends on your department and team needs, typically ranging from weekly to monthly. These in-person moments help us build stronger connections, spark new ideas, and celebrate our successes together. Most days, you can choose what works best for you, while staying in tune with your team's goals.
 
We're excited to welcome you to our team and look forward to achieving great things together - both in the office and wherever you work best!
 
The Most Important Thing
 
We're looking for candidates that can provide examples of how they have demonstrated Magnet CODE in their previous experiences:
 
CARE - We care about each other and our mission to make a difference in the world.
OWN - We are accountable for our results - while never forgetting to act with integrity, empathy, and respect.
DEDICATE - We put our heart and soul into meeting the needs of our customers and helping them serve the people they protect.
EVOLVE - We are constantly innovating and exploring new ways to work together to make an impact with our work.
 
Here at Magnet Forensics, we are committed to continuous learning and are focused on building a diverse and inclusive workforce. This commitment will be reflected in our hiring processes and embedded in our values and how we treat one another. If you're interested in this role, but do not meet all of the qualifications listed above, we encourage you to apply anyways.
 
Magnet Forensics is an Equal Opportunity Employer and considers applicants for employment without regard to race, colour, religion, sex, orientation, national origin, age, disability, genetics or any other basis forbidden under federal, provincial, or local law. We are committed to providing an inclusive, accessible recruitment process and work environment. Accommodation is available to all applicants upon request throughout the hiring process. Please contact [email protected] should you require any accommodations.
 
All offers of employment at Magnet are contingent upon satisfactory completion of a background check. All background checks will be conducted in accordance with all applicable laws. Magnet will consider each position's job duties, among other factors, in determining what constitutes satisfactory completion of the background check. Refusal to consent to a background check may be grounds for revoking an offer of employment.
 
US Applicants: Magnet Forensics participates in E-Verify and will provide the federal government with your Form I-9 information to confirm that you are authorized to work in the U.S.
 
Magnet Forensics handles and uses personal data of job applicants in line with its Recruitment Privacy Policy found here. 
Magnet does not use artificial intelligence (defined as a machine-based system that infers from input to generate outputs such as predictions, content, recommendations, or decisions) for screening, assessing, or selecting applicants. Should this practice change, we will update this disclosure accordingly.
apply for this job