1

Devops Engineer Site Reliability Engineer Jobs in Vaughan, ON

Senior Site Reliability Engineer I

Toronto, ON ยท On-site

CA$153K - CA$277K/yr

Professional Experience: 5+ years of experience as a DevOps, or Site Reliability Engineer in a high-scale production environment * NGINX Mastery: Deep hands-on experience configuring, troubleshooting ...

As a DevOps Engineer, you will be responsible for designing, building, and operating automated ... (SRE mindset) * Implement and maintain: * Monitoring * Logging * Alerting * Ensure end-to-end ...

Skills and Qualifications: * 3+ years working in System Engineering, DevOps, Site Reliability Engineering, or similar capacity preferably managing SaaS environments. * 3+ years of experience working ...

DevOps Engineer

Toronto, ON ยท On-site

CA$90K - CA$115K/yr

Degree in Software Engineering, IT, or related discipline, or equivalent practical experience * 5+ years of experience in DevOps/SRE/system administration or related internships/projects

DevOps Engineer

Toronto, ON ยท On-site

CA$90K - CA$115K/yr

Degree in Software Engineering, IT, or related discipline, or equivalent practical experience * 5+ years of experience in DevOps/SRE/system administration or related internships/projects

WhatWe'reLooking For * 5+ years of relevant experience in DevOps, Site Reliability Engineering, or cloud operations. * Extensive experience with Azure services, Azure DevOps pipelines, and automation ...

Site Reliability Engineer Lead

Toronto, ON ยท On-site

CA$125K - CA$148K/yr

Strong understanding of cloud infrastructure, database platforms, DevOps, automation, and platform ... or site reliability practices. We recognize that strong candidates may bring relevant experience ...

DevOps

Toronto, ON ยท On-site

Who You Are * 5+ years of hands-on experience in DevOps, Site Reliability Engineering, Platform Engineering, Cloud Infrastructure, or a similar engineering role, particularly within the Microsoft ...

Establishes reliability standards, service level objectives, and automation practices across ... Azure certification (e.g., Azure Solutions Architect Expert or DevOps Engineer Expert). Experience ...

Showing results 41-60

Devops Engineer Site Reliability Engineer information

What is a DevOps engineer site reliability engineer?

DevOps Engineers and Site Reliability Engineers (SREs) are IT professionals who focus on optimizing the development, deployment, and reliability of software systems. DevOps Engineers bridge the gap between development and operations by automating processes, managing CI/CD pipelines, and collaborating across teams. SREs, on the other hand, apply software engineering principles to system administration tasks, with a strong emphasis on reliability, scalability, and performance. Both roles aim to improve the efficiency, stability, and consistency of software delivery, but SREs typically place a heavier focus on system reliability and uptime.

How do DevOps engineers site reliability engineers typically collaborate with development and operations teams?

DevOps Engineers and Site Reliability Engineers (SREs) work closely with both development and operations teams to ensure smooth software delivery and system reliability. They often act as bridges, facilitating communication and implementing automation tools, CI/CD pipelines, and monitoring solutions. Collaboration can include participating in sprint planning, incident response, and root cause analysis meetings. They also play a key role in educating team members on best practices for infrastructure as code, deployment processes, and system scalability. This collaborative approach helps to foster a culture of shared responsibility for the software lifecycle.

What are the key skills and qualifications needed to thrive as a DevOps engineer site reliability engineer?

To thrive as a DevOps Engineer or Site Reliability Engineer, you need strong skills in systems administration, automation, cloud platforms, and scripting, typically supported by a degree in computer science or related field. Familiarity with tools such as Docker, Kubernetes, Jenkins, Terraform, and monitoring solutions, as well as certifications like AWS Certified DevOps Engineer or Google Professional SRE, is highly valued. Collaboration, problem-solving, and effective communication are crucial soft skills for working with cross-functional teams and responding to incidents. These skills are essential for maintaining scalable, reliable, and secure infrastructure while enabling rapid development and deployment cycles.

What is the difference between Devops Engineer Site Reliability Engineer vs Site Reliability Engineer?

AspectDevops EngineerSite Reliability Engineer
Primary FocusAutomating deployment, CI/CD pipelines, and infrastructure managementEnsuring system reliability, scalability, and incident response
Skills & CertificationsCloud platforms, scripting, automation tools, CI/CDMonitoring, incident management, system architecture
Work EnvironmentDevelopment and operations teams, often in cloud environmentsOperations teams, on-call duties, production systems
Industry UsageTech companies, startups, cloud service providersLarge-scale tech firms, cloud providers, enterprises

While both roles overlap in infrastructure and automation, Devops Engineers focus on streamlining deployment processes, whereas Site Reliability Engineers prioritize system stability and incident management. Understanding these distinctions helps in choosing the right career path or job role.

What cities near Vaughan, ON are hiring for Devops Engineer Site Reliability Engineer jobs?

Cities near Vaughan, ON with the most Devops Engineer Site Reliability Engineer job openings:

Infographic showing various Devops Engineer Site Reliability Engineer job openings in Vaughan, ON as of July 2026, with employment types broken down into 1% As Needed, 78% Full Time, 17% Part Time, 1% Temporary, and 3% Contract. Highlights an 94% Physical, 2% Hybrid, and 4% Remote job distribution.

Senior Site Reliability Engineer I

Toronto, ON โ€ข On-site

Braze
Software Developmentย โ€ขย 1 - 5K employees

CA$153K - CA$277K/yr

Full-time

Posted 15 days ago


Job description

Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing services and platforms running smoothly, ensuring consistent system reliability and maximum infrastructure uptime. SREs blend the roles of sensible system administrators and software engineers, applying sound engineering principles, operational discipline, and mature automation to our production environments. We specialize in distributed systems, operating at an extraordinary scale that includes serving over 3.3 billion monthly active users, collecting hundreds of billions of data points each month, and delivering billions of messages to our customers' end-users daily.

This Senior SRE role is specifically focused on supporting the Braze Ruby on Rails monolith and our growing fleet of Go API services. This position is embedded with product engineering teams who own APIs, taking their high-level feature requirements and helping translate them into reliable, highly scalable technology stacks. You will be responsible for the ingress fleets and API ingestion layers which serve as the main entry point for how the entire internet talks to Braze. Because of the highly operational and foundational nature of these ingress fleets, a deep, hands-on experience of both NGINX and Kubernetes is important.

WHAT YOU'LL DO

Lead NGINX & Kubernetes Ingress Infrastructure

  • Architect and Operate Ingress Fleets: Configure, tune, and operate Braze's high-performance NGINX routing, proxying, and ingress controller layers that manage massive, real-time API ingestion traffic
  • Own and Expand Scaling Routines: Own, expand, and tune automated scaling routines for our high-throughput, API-based services, leveraging RED (Request, Error, Duration) metrics, Horizontal Pod Autoscalers (HPA), and customized scaling policies to ensure seamless traffic handling under extreme spikes

Partner with Product Teams on Resilient Architectures

  • Translate Product Requirements: Partner directly with product engineering teams owning API features to translate product requirements into resilient, highly available, and scalable technology stacks
  • Manage SLIs, SLOs, and Error Budgets: Establish meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for API services, helping teams navigate error budgets to balance rapid feature deployment with stability
  • Systems Design & Capacity Planning: Conduct in-depth systems design, bottleneck profiling, and capacity planning to ensure Braze meets its strict enterprise-grade SLAs

Incident Response & On-Call Excellence

  • Proactive On-Call Rotation: Participate in a PagerDuty on-call rotation, utilizing shifts not just to resolve structural alerts, but to update runbooks and proactively prevent incidents from recurring
  • Blameless Retrospectives: Lead root-cause analysis (RCA) and blameless retrospectives for availability and performance incidents, translating operational learnings into permanent system improvements

WHO YOU ARE

Required Experience & Skills:

  • Professional Experience: 5+ years of experience as a DevOps, or Site Reliability Engineer in a high-scale production environment
  • NGINX Mastery: Deep hands-on experience configuring, troubleshooting, and operating high-performance NGINX proxying, routing, and ingress controller layers under heavy traffic loads
  • Kubernetes Orchestration: In-depth, hands-on proficiency with Kubernetes administration, cluster networking, container orchestration, cluster scheduling, and container deployment
  • Linux Fundamentals: Excellent, OS-level understanding of Linux/Unix internals, including disk I/O, memory allocation, TCP/IP networking, and process management
  • Programming Proficiency: Strong programming/scripting skills-Ruby and/or Go preferred (or equivalent languages like Python, Java) to build custom automated tools, scripting, and platform frameworks
  • Infrastructure as Code: Experience managing infrastructure using IaC technologies such as Terraform, Ansible, Chef, or similar frameworks
  • Systems Thinking: Strong conceptual understanding of systems design-interfaces, boundaries, failure modes, edge cases, and cascading effects across distributed architectures
  • Collaborative Mindset: Excellent documentation practices and a high comfort level collaborating asynchronously across remote-first, global engineering teams

Preferred / Nice to Have:

  • Data Technologies: Familiarity with data-tier architectures in our stack, including Redis, Kafka, Postgres, or MongoDB
  • Observability Systems: Experience using monitoring, profiling, and observability platforms (e.g., Prometheus, Grafana, Datadog or similar tools) to narrow in on incidents quickly
  • Cloud Infrastructures: Practical experience deploying and managing infrastructure services in major cloud environments (AWS, GCP, Azure)

For candidates based in Canada, the pay range at the start of employment for this position is expected to be between CA$153,815 - CA$277,000/year, with an expected On Target Earnings (OTE) between CA$172,000 - CA$308,400/year (including performance-based or variable compensation (bonus or commission). Your particular offer may vary depending on multiple individual factors, including market location, job-related knowledge, skills, and experience. In addition to cash compensation, this role qualifies for a comprehensive Total Rewards package that includes equity grants of restricted stock (RSUs) so that you will own a piece of our company.