1

Process Reliability Manager Jobs in Ontario (NOW HIRING)

Site Reliability Engineer

Toronto, ON ยท On-site +1

CA$125K - CA$250K/yr

Improve provisioning, configuration management, testing, and deployment automation * Help plan ... We may use artificial intelligence (AI) tools to support parts of the hiring process, such as ...

Site Reliability Engineer

Toronto, ON ยท Hybrid

CA$100K - CA$125K/yr

... supplier management, tax compliance, and treasury. Tipalti partners with leading financial ... Driving incident response and post-mortem processes, fostering a culture of continuous improvement.

... management processes Execute disaster recovery, configuration management, and infrastructure readiness tasks Keeping oversight and coordinating routine maintenance, deployments, refreshes, rollbacks ...

... process. You'll engage in and often lead architectural discussions, reduce toil, and deliver ... You'll be a key voice in observability, change management, and service scalability, providing ...

Showing results 41-60

Process Reliability Manager information

What is a process reliability manager?

A Process Reliability Manager is a professional responsible for ensuring that manufacturing or production processes operate efficiently, consistently, and with minimal downtime. They analyze process data, identify areas for improvement, and implement strategies to enhance equipment reliability and overall process performance. By collaborating with maintenance, engineering, and operations teams, they help reduce failures, optimize productivity, and maintain quality standards. Their work is crucial for minimizing costs and ensuring that production targets are met safely and reliably.

What are the key skills and qualifications needed to thrive as a process reliability manager, and why are they important?

To thrive as a Process Reliability Manager, you need a strong background in engineering, process optimization, and reliability analysis, often supported by a degree in engineering and experience in manufacturing or industrial settings. Familiarity with reliability-centered maintenance (RCM), root cause analysis tools, and data analysis software such as SAP or Maximo is typically required. Exceptional problem-solving, leadership, and communication skills help drive cross-functional teams and foster a culture of continuous improvement. These skills are crucial to ensure equipment reliability, minimize downtime, and optimize operational efficiency within complex production environments.

How does a process reliability manager typically collaborate with maintenance and production teams to achieve operational goals?

A Process Reliability Manager works closely with both maintenance and production teams to identify areas of improvement in equipment reliability and process efficiency. This often involves facilitating cross-functional meetings, analyzing downtime data, and implementing preventive maintenance strategies. Clear communication and teamwork are key, as the role requires aligning the objectives of different departments to minimize unplanned outages and optimize production output. By fostering a proactive culture and sharing best practices, the Process Reliability Manager helps ensure the plant operates smoothly and efficiently.

What is the difference between Process Reliability Manager vs Maintenance Engineer?

AspectProcess Reliability ManagerMaintenance Engineer
CertificationsReliability certifications, Six Sigma, PMPMechanical/Electrical certifications, HVAC, PLC certifications
Work EnvironmentManufacturing plants, industrial facilitiesFactories, equipment maintenance sites
Industry UsageFocus on reliability, uptime, and process optimizationFocus on equipment repair, preventive maintenance

The Process Reliability Manager primarily focuses on improving equipment reliability and process efficiency through data analysis and strategic planning. In contrast, Maintenance Engineers handle the hands-on repair and maintenance of machinery. Both roles are essential in manufacturing, but the Process Reliability Manager emphasizes proactive reliability strategies, while Maintenance Engineers focus on reactive and preventive maintenance tasks.

What cities in Ontario are hiring for Process Reliability Manager jobs?

Cities in Ontario with the most Process Reliability Manager job openings:

Site Reliability Engineer

Boson AI

Toronto, ON โ€ข On-site, Remote

CA$125K - CA$250K/yr

Full-time

Re-posted 9 days ago


Job description

About The Role
ย 
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work.
ย 
Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from "it works" to dependable, observable, and scalable.
ย 
You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area-networking, cluster scheduling, storage, GPU systems, or AI infrastructure- and the curiosity and judgment to collaborate across the rest.
ย 
Responsibilities
  • Design, operate, and improve reliable infrastructure for AI training and inference workloads
  • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms
  • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate
  • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads
  • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements
  • Improve provisioning, configuration management, testing, and deployment automation
  • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management
  • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards
Minimum Qualifications
  • 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role
  • Strong hands-on expertise in at least one of the following:
    • Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand
    • Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms
    • Distributed storage, particularly Ceph
    • GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting
    • AI training or model-serving infrastructure
  • Experience operating production systems with a focus on availability, performance, security, and automation
  • Strong Linux administration and scripting skills
  • A systematic approach to troubleshooting across multiple layers of a complex system
  • Clear written and verbal communication skills, including the ability to work effectively with a distributed team
Preferred Qualifications
  • Experience supporting GPU-intensive AI or HPC environments
  • Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet
  • Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling
  • Experience operating or tuning Ceph clusters
  • Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems
  • Experience with hardware provisioning, firmware management, and bare-metal automation
  • Experience running large-scale distributed training or high-throughput inference workloads
  • Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure
$125,000 - $250,000 a year
Boson AI is building AI systems for real-world, business-critical use. If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we'd love to hear from you.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
apply for this job