2

Remote Site Reliability Engineer Intern Jobs in Toronto, ON

Site Reliability Engineer

Toronto, ON ยท On-site +1

CA$125K - CA$250K/yr

Based in Toronto or remote, you will work across the systems that enable large-scale AI training ... site reliability engineering, infrastructure engineering, systems engineering, or a related ...

To do that we are eager to add a highly skilled DevOps / SRE Engineer Engineer to our incredible ... However, we will consider remote applicants +/- 3 hours from eastern time zone. Why work here * We ...

To do that we are eager to add a highly skilled DevOps / SRE Engineer Engineer to our incredible ... However, we will consider remote applicants +/- 3 hours from eastern time zone. Why work here * We ...

Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience) Type: Contract Compensation: $85/hour Location: Remote Role Responsibilities * Use frontier AI coding agents to complete and evaluate ...

Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience) Type: Contract Compensation: $85/hour Location: Remote Role Responsibilities * Use frontier AI coding agents to complete and evaluate ...

You will operate with a high degree of autonomy, applying SRE principles and an engineering-first mindset to ensure the reliability, performance, and availability of mission-critical healthcare ...

You will operate with a high degree of autonomy, applying SRE principles and an engineering-first mindset to ensure the reliability, performance, and availability of mission-critical healthcare ...

Design and own the SRE function for Level 1 data ingestion across all GCP deployments: alert policy ... Comfort with distributed teams and experience managing offshore or remote team members * Ability to ...

Senior Platform Engineer

Toronto, ON ยท On-site +1

CA$100K - CA$150K/yr

The SPS SRE team is responsible for delivering highly available platform services and deployment ... Location: This role follows a remote work model for candidates who are based in Canada. What We ...

DevOps Developer

Toronto, ON ยท On-site +1

CA$125K - CA$140K/yr

This is a remote location open to candidates legally authorized to work in Canada. What you will be ... SRE, or software engineering roles with strong operational ownership. * Strong TypeScript ...

next page

Showing results 1-20

Remote Site Reliability Engineer Intern information

What is the difference between Remote Site Reliability Engineer Intern vs Remote DevOps Engineer Intern?

AspectRemote Site Reliability Engineer InternRemote DevOps Engineer Intern
FocusEnsuring system reliability, uptime, and performanceAutomating deployment, integration, and infrastructure management
SkillsMonitoring, incident response, scripting, cloud platformsCI/CD, automation tools, scripting, cloud services
Work EnvironmentCollaborates with SRE and operations teamsWorks with development and operations teams
CertificationsOften preferred: Linux, cloud certificationsOften preferred: Linux, cloud, and automation certifications

While both roles involve cloud and scripting skills, the Remote Site Reliability Engineer Intern focuses on system stability and uptime, whereas the Remote DevOps Engineer Intern emphasizes automation and deployment processes. Understanding these differences helps candidates target their skills and career goals effectively.

What are popular job titles related to Remote Site Reliability Engineer Intern jobs in Toronto, ON? For Remote Site Reliability Engineer Intern jobs in Toronto, ON, the most frequently searched job titles are:
What job categories do people searching Remote Site Reliability Engineer Intern jobs in Toronto, ON look for? The top searched job categories for Remote Site Reliability Engineer Intern jobs in Toronto, ON are:
Infographic showing various Remote Site Reliability Engineer Intern job openings in Toronto, ON as of August 2026, with employment types broken down into 68% Full Time, 19% Part Time, and 13% Contract. Highlights an 100% Remote job distribution.

Site Reliability Engineer

Boson AI

Toronto, ON โ€ข On-site, Remote

CA$125K - CA$250K/yr

Full-time

Posted 27 days ago


Job description

About The Role
ย 
Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work.
ย 
Based in Toronto or remote, you will work across the systems that enable large-scale AI training and serving: high-performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking complex infrastructure from "it works" to dependable, observable, and scalable.
ย 
You do not need to be an expert in every layer of the stack. We are looking for deep strength in at least one area-networking, cluster scheduling, storage, GPU systems, or AI infrastructure- and the curiosity and judgment to collaborate across the rest.
ย 
Responsibilities
  • Design, operate, and improve reliable infrastructure for AI training and inference workloads
  • Own and automate operational workflows across one or more core areas: networking, compute allocation, storage, GPU/server configuration, or AI platforms
  • Build monitoring, alerting, runbooks, and incident-response practices that make systems easier to operate
  • Diagnose performance, capacity, and reliability issues across hardware, operating systems, networks, schedulers, and distributed workloads
  • Partner closely with ML, research, and platform teams to translate workload needs into practical infrastructure improvements
  • Improve provisioning, configuration management, testing, and deployment automation
  • Help plan cluster growth, capacity allocation, upgrades, and lifecycle management
  • Contribute to a thoughtful reliability culture through documentation, post-incident learning, and pragmatic engineering standards
Minimum Qualifications
  • 4+ years of experience in site reliability engineering, infrastructure engineering, systems engineering, or a related production-operations role
  • Strong hands-on expertise in at least one of the following:
    • Networking, including firewalls, switching, routing, ASN/BGP configuration, or InfiniBand
    • Cluster and systems allocation with Kubernetes, SLURM, MAAS, or similar platforms
    • Distributed storage, particularly Ceph
    • GPU and server administration, including CUDA drivers, firmware, BIOS, and hardware troubleshooting
    • AI training or model-serving infrastructure
  • Experience operating production systems with a focus on availability, performance, security, and automation
  • Strong Linux administration and scripting skills
  • A systematic approach to troubleshooting across multiple layers of a complex system
  • Clear written and verbal communication skills, including the ability to work effectively with a distributed team
Preferred Qualifications
  • Experience supporting GPU-intensive AI or HPC environments
  • Experience with NVIDIA GPUs, CUDA, NCCL, and high-performance interconnects - Experience with InfiniBand, RDMA, RoCE, or 100Gb+ Ethernet
  • Familiarity with Kubernetes, SLURM, MAAS, Terraform, Ansible, or similar infrastructure tooling
  • Experience operating or tuning Ceph clusters
  • Familiarity with observability tooling such as Prometheus, Grafana, and centralized logging systems
  • Experience with hardware provisioning, firmware management, and bare-metal automation
  • Experience running large-scale distributed training or high-throughput inference workloads
  • Familiarity with cloud and hybrid infrastructure across AWS, GCP, or Azure
$125,000 - $250,000 a year
Boson AI is building AI systems for real-world, business-critical use. If you enjoy solving difficult infrastructure problems and want your work to directly enable the next generation of AI products, we'd love to hear from you.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
apply for this job