1

Network Reliability Engineer Jobs in Toronto, ON

Site Reliability Engineer

Toronto, ON · On-site +1

CA$125K - CA$250K/yr

We are looking for deep strength in at least one area-networking, cluster scheduling, storage, GPU ... site reliability engineering, infrastructure engineering, systems engineering, or a related ...

Site Reliability Engineer (.Net)

Toronto, ON · Hybrid

CA$100K - CA$125K/yr

As a Site Reliability Engineer, you will play a crucial role in enhancing the reliability ... Basic understanding of some of the lower levels of software frameworks and networking.

WHY THIS ROLE IS IMPORTANT TO US As a Senior Site Reliability Engineer, you will be embedded within ... networking, virtualization, containerization (Kubernetes, Docker) Comfort with Linux and Windows ...

We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ... Confident navigating Linux environments, checking system logs, and analyzing network traffic.

You will operate with a high degree of autonomy, applying SRE principles and an engineering-first ... Manage and optimize CiscoThousandEyesfor network path visibility, internet performance monitoring ...

You will operate with a high degree of autonomy, applying SRE principles and an engineering-first ... Manage and optimize CiscoThousandEyesfor network path visibility, internet performance monitoring ...

Lead Network and Cloud Architect

Mississauga, ON · On-site +1

CA$122K - CA$162K/yr

Own network capacity planning, including optimization of performance, scalability, and cost ... Drive innovation through infrastructure automation, orchestration, and reliability engineering ...

next page

Showing results 1-20

Network Reliability Engineer information

See Toronto, ON salary details

$41K

$121.1K

$167K

How much do network reliability engineer jobs pay per year?

As of Aug 7, 2026, the average yearly pay for network reliability engineer in Toronto, ON is $121,123.00, according to ZipRecruiter salary data. Most workers in this role earn between $102,591.00 and $144,105.00 per year, depending on experience, location, and employer.

What is a network reliability engineer?

A Network Reliability Engineer (NRE) is an IT professional responsible for ensuring the reliability, performance, and scalability of network systems. They combine skills in networking, software engineering, and automation to proactively detect and resolve potential network issues before they affect users. NREs often design and implement monitoring tools, automate network management tasks, and work to improve the overall stability of network infrastructure. Their goal is to minimize downtime and ensure seamless connectivity across an organization’s network.

What is the difference between Network Reliability Engineer vs Network Operations Center (NOC) Technician?

AspectNetwork Reliability EngineerNetwork Operations Center (NOC) Technician
CertificationsCCNA, CCNP, Network+CCNA, Network+
Work EnvironmentDesign, analyze, and improve network infrastructureMonitor, troubleshoot, and maintain networks in real-time
Employer & Industry UsageTelecom, large enterprises, cloud providersISPs, data centers, enterprise networks
Common Search & ComparisonFocus on network reliability and designFocus on network monitoring and incident response

The main difference is that Network Reliability Engineers focus on designing and improving network systems to ensure long-term reliability, while NOC Technicians monitor and troubleshoot networks in real-time to resolve issues quickly. Both roles require relevant certifications and are essential in maintaining network performance, but they serve different functions within network management.

What are the key skills and qualifications needed to thrive as a network reliability engineer, and why are they important?

To thrive as a Network Reliability Engineer, you need a strong background in computer networking, network protocols, troubleshooting, and often a degree in computer science or a related field. Familiarity with tools like Wireshark, Nagios, Cisco IOS, and certifications such as CCNA or CCNP are commonly required. Analytical thinking, proactive problem-solving, and effective communication are standout soft skills in this role. These skills are crucial to maintaining reliable network operations, minimizing downtime, and ensuring seamless communication across organizational systems.

What are some common challenges faced by network reliability engineers, and how are they typically addressed?

Network Reliability Engineers often encounter challenges such as diagnosing intermittent connectivity issues, managing network upgrades with minimal downtime, and maintaining high availability during peak traffic. These are typically addressed by leveraging robust monitoring tools, implementing automation for routine tasks, and collaborating closely with software engineers, network administrators, and incident response teams. Staying current with evolving network technologies and best practices is essential for effectively identifying and resolving problems before they impact users.
What are popular job titles related to Network Reliability Engineer jobs in Toronto, ON? For Network Reliability Engineer jobs in Toronto, ON, the most frequently searched job titles are:
What job categories do people searching Network Reliability Engineer jobs in Toronto, ON look for? The top searched job categories for Network Reliability Engineer jobs in Toronto, ON are:
Infographic showing various Network Reliability Engineer job openings in Toronto, ON as of June 2026, with employment types broken down into 75% Full Time, and 25% Contract. Highlights an 100% In-person job distribution, with an average salary of $121,123 per year, or $58.2 per hour.

Site Reliability Engineer

Boson AI

Toronto, ON • On-site, Remote

CA$125K - CA$250K/yr

Full-time

Posted 24 days ago


Job description

About The Role
 
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work.
 
Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from "it works" to dependable, observable, and scalable.
 
You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area-networking, cluster scheduling, storage, GPU systems, or AI infrastructure- and the curiosity and judgment to collaborate across the rest.
 
Responsibilities
  • Design, operate, and improve reliable infrastructure for AI training and inference workloads
  • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms
  • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate
  • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads
  • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements
  • Improve provisioning, configuration management, testing, and deployment automation
  • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management
  • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards
Minimum Qualifications
  • 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role
  • Strong hands-on expertise in at least one of the following:
    • Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand
    • Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms
    • Distributed storage, particularly Ceph
    • GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting
    • AI training or model-serving infrastructure
  • Experience operating production systems with a focus on availability, performance, security, and automation
  • Strong Linux administration and scripting skills
  • A systematic approach to troubleshooting across multiple layers of a complex system
  • Clear written and verbal communication skills, including the ability to work effectively with a distributed team
Preferred Qualifications
  • Experience supporting GPU-intensive AI or HPC environments
  • Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet
  • Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling
  • Experience operating or tuning Ceph clusters
  • Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems
  • Experience with hardware provisioning, firmware management, and bare-metal automation
  • Experience running large-scale distributed training or high-throughput inference workloads
  • Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure
$125,000 - $250,000 a year
Boson AI is building AI systems for real-world, business-critical use. If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we'd love to hear from you.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
apply for this job