1

Reliability Manager Jobs in Markham, ON (NOW HIRING)

Site Reliability Engineer

Toronto, ON ยท On-site +1

CA$125K - CA$250K/yr

We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind ... Improve provisioning, configuration management, testing, and deployment automation * Help plan ...

SRE is part of a global organization that leverages the latest technology to communicate with our ... You'll be a key voice in observability, change management, and service scalability, providing ...

Site Reliability Engineer (.Net)

Toronto, ON ยท Hybrid

CA$100K - CA$125K/yr

... supplier management, tax compliance, and treasury. Tipalti partners with leading financial ... Lead reliability-focused practices such as SLO (Service Level Objective) design and implementation ...

... Manage infrastructure deployment pipelines and troubleshoot onboarding and operational issues ... in Site Reliability, DevOps, or Cloud Engineering roles Expertise with Microsoft Azure; AWS ...

Showing results 21-40

Reliability Manager information

See Markham, ON salary details

$74.4K

$107.2K

$136.9K

How much do reliability manager jobs pay per year?

As of Aug 10, 2026, the average yearly pay for reliability manager in Markham, ON is $107,248.00, according to ZipRecruiter salary data. Most workers in this role earn between $93,791.00 and $117,476.00 per year, depending on experience, location, and employer.

What is the role of a reliability manager?

A reliability manager oversees the maintenance and reliability of equipment and systems within an organization to ensure optimal performance and minimize downtime. They analyze failure data, develop maintenance strategies, and implement continuous improvement processes, often using tools like root cause analysis and reliability-centered maintenance. Strong analytical skills, technical knowledge, and certifications such as Certified Reliability Engineer (CRE) are typically required for this role.

Is reliability engineering in demand?

Reliability engineering is in high demand across industries such as manufacturing, energy, and aerospace, as companies prioritize system uptime and maintenance efficiency. Reliability Managers with skills in data analysis, failure modes, and certifications like RCM are sought after to improve equipment performance and reduce downtime.

What are the key skills and qualifications needed to thrive as a reliability manager?

A Reliability Manager needs strong analytical skills, a solid background in engineering or maintenance, and experience with reliability-centered maintenance methodologies. Familiarity with tools like Failure Mode and Effects Analysis (FMEA), Root Cause Analysis (RCA), and certifications such as Certified Reliability Engineer (CRE) are often required. Leadership, problem-solving, and the ability to communicate complex technical information clearly are crucial soft skills for this role. These skills help ensure equipment uptime, optimize maintenance processes, and foster a culture of continuous improvement within the organization.

What does a reliability manager do?

A Reliability Manager is responsible for ensuring that equipment, processes, and systems operate efficiently and consistently to minimize downtime and maximize performance. They develop and implement reliability strategies, conduct root cause analyses, and oversee preventive and predictive maintenance programs. Their role involves working closely with maintenance teams, engineers, and production staff to improve asset reliability and extend equipment lifespan. Additionally, they analyze failure data, recommend improvements, and help optimize operational costs through reliability-centered maintenance practices.

What cities near Markham, ON are hiring for Reliability Manager jobs? Cities near Markham, ON with the most Reliability Manager job openings:
Infographic showing various Reliability Manager job openings in Markham, ON as of August 2026, with employment types broken down into 88% Full Time, 11% Part Time, and 1% Contract. Highlights an 84% Physical, 3% Hybrid, and 13% Remote job distribution, with an average salary of $107,248 per year, or $51.6 per hour.

Site Reliability Engineer

Boson AI

Toronto, ON โ€ข On-site, Remote

CA$125K - CA$250K/yr

Full-time

Posted 28 days ago


Job description

About The Role
ย 
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work.
ย 
Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from "it works" to dependable, observable, and scalable.
ย 
You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area-networking, cluster scheduling, storage, GPU systems, or AI infrastructure- and the curiosity and judgment to collaborate across the rest.
ย 
Responsibilities
  • Design, operate, and improve reliable infrastructure for AI training and inference workloads
  • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms
  • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate
  • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads
  • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements
  • Improve provisioning, configuration management, testing, and deployment automation
  • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management
  • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards
Minimum Qualifications
  • 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role
  • Strong hands-on expertise in at least one of the following:
    • Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand
    • Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms
    • Distributed storage, particularly Ceph
    • GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting
    • AI training or model-serving infrastructure
  • Experience operating production systems with a focus on availability, performance, security, and automation
  • Strong Linux administration and scripting skills
  • A systematic approach to troubleshooting across multiple layers of a complex system
  • Clear written and verbal communication skills, including the ability to work effectively with a distributed team
Preferred Qualifications
  • Experience supporting GPU-intensive AI or HPC environments
  • Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet
  • Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling
  • Experience operating or tuning Ceph clusters
  • Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems
  • Experience with hardware provisioning, firmware management, and bare-metal automation
  • Experience running large-scale distributed training or high-throughput inference workloads
  • Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure
$125,000 - $250,000 a year
Boson AI is building AI systems for real-world, business-critical use. If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we'd love to hear from you.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
apply for this job