1

Site Reliability Engineer Manager Jobs in Toronto, ON

We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability ... Create, manage, and support pipelines that the application support teams will be utilizing to ...

The OPS Site Reliability Engineer will be a focal role owning and ensuring the fluent operations of Managed Services offerings in the KPMG production cloud environment. The role will be focusing on ...

Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing services ... Manage SLIs, SLOs, and Error Budgets: Establish meaningful Service Level Indicators (SLIs) and ...

Showing results 21-40

Site Reliability Engineer Manager information

See Toronto, ON salary details

$124.5K

$155.9K

$187.5K

How much do site reliability engineer manager jobs pay per year?

As of Sep 3, 2026, the average yearly pay for site reliability engineer manager in Toronto, ON is $155,875.00, according to ZipRecruiter salary data. Most workers in this role earn between $146,014.00 and $165,578.00 per year, depending on experience, location, and employer.

What is a site reliability engineer manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What are the key skills and qualifications needed to thrive as a site reliability engineer manager?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.

How does a site reliability engineer manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How much do site reliability engineer managers get paid?

Site Reliability Engineer Managers typically earn between $120,000 and $180,000 annually, depending on experience, location, and company size. They often oversee teams responsible for system reliability, incident response, and infrastructure automation, requiring strong leadership and technical skills.

Is a Site Reliability Engineer Manager a stressful job?

A Site Reliability Engineer Manager role can be stressful due to the responsibility of maintaining system uptime, managing incident responses, and ensuring reliability across complex infrastructure. The job often involves working under pressure, handling outages, and coordinating teams, but it also offers opportunities for problem-solving and leadership. Stress levels vary depending on company size, team structure, and workload management skills.

What are the most commonly searched types of Site Reliability Engineer jobs in Toronto, ON?

The most popular types of Site Reliability Engineer jobs in Toronto, ON are:

Infographic showing various Site Reliability Engineer Manager job openings in Toronto, ON as of August 2026, with employment types broken down into 88% Full Time, 11% Part Time, and 1% Contract. Highlights an 80% Physical, 4% Hybrid, and 16% Remote job distribution, with an average salary of $155,875 per year, or $74.9 per hour.

Senior Director, Head, SRE and Production Operations

Royal Bank of Canada

Toronto, ON • On-site

Full-time

Posted 21 days ago


Job description

Job Description

What is the Opportunity?

As Head of SRE and Production Operations, you will lead the vision, design, development, implementation, and support of Site Reliability Engineering (SRE) solutions for all applications supported by Commercial & Payments Technology (CPT) within Technology & Operations (T&O). This strategic role is instrumental in establishing world-class operational excellence, enabling automated end-to-end observability, and driving continuous reliability improvements across critical enterprise applications.

You will work collaboratively with teams across multiple lines of business and organizational partners to establish SRE as a key enabler of business success. As a strategic leader with deep technical expertise, you will drive execution excellence while building and developing high-performing teams capable of managing complex reliability and production support challenges at enterprise scale.

What Will You Do?

Technical Leadership

  • Set the vision for SRE product offerings including monitoring, alerting, machine learning anomaly detection, self-healing capabilities, and reliability testing

  • Drive best-in-class technical solutions by tracking industry-leading practices and applying them to the RBC environment

  • Lead thought leadership and out-of-the-box thinking to ensure SRE tools and processes align with strategic objectives

  • Run engineering practice forums to facilitate the review of SRE solutions and maintain alignment with team and enterprise vision

  • Enable automated end-to-end observability across all critical applications within Commercial & Payments Technology (CPT)

  • Leverage unit, department, and enterprise-wide teams to develop better solutions and achieve a cross-enterprise mindset

Production Support

  • Perform production support role, including off-hours support responsibilities

  • Assist in incident management and problem management for applications in scope

  • Continuously evaluate what went well, what went wrong, and what can be done to improve and prevent issues in the future

  • Maintain technology currency through server patching, certificate renewal, and other maintenance activities with a keen eye on automating opportunities

  • Ensure availability and uptime of applications in scope according to defined service level objectives

  • Ensure compliance of all systems and applications in scope, including maintaining segregation of duties

  • Leverage AI and other key technologies to drive efficiencies and meet Commercial & Payments Technology (CPT), Technology & Operations (T&O) and respective business goals and KPIs

Strategy

  • Drive the overall SRE strategy within Innovation, Intelligent Operations, and partner groups, owning roadmap development

  • Lead the team through execution of the SRE roadmap for Innovation and partner groups

  • Drive participation in strategic steering groups

  • Lead the adoption of new technologies within SRE

  • Enable teams to meet objectives by implementing changes to processes, tools, and methods that result in increased agility and effective cost management

  • Develop governance of the unit in collaboration with executive leadership to provide structure and definition for effectively managing the Innovation SRE Organization

  • Drive industry best standards and practices for the most effective and efficient SRE and production support services

People Leadership

  • Provide expertise, direction, coaching, and development to build team capability and succession planning

  • Select and build a high-performing, diverse team that leverages individual capabilities and strengths

  • Promote a mindset for sustained success, growth, diversity, and an overall engineering mindset

  • Ensure that employees understand RBC's vision and support and reinforce targeted behaviours that contribute to RBC goals

  • Provide focus and clarity in establishing individual goals, driving performance enablement, supporting career development, and rewarding strong performance

  • Attract, hire, and retain top talent

  • Manage key vendor relationships, contractors, and respective performance management to enable successful execution of SRE and production support services

What do you need to succeed?

Must-have

  • Minimum 8-12 years of experience leading SRE, production operations, or infrastructure engineering teams at scale

  • Deep technical expertise in observability, monitoring, and incident management platforms

  • Demonstrated experience with cloud platforms, containerization, and automation technologies

  • Strong understanding of reliability engineering principles, SLOs, SLIs, and error budgets

  • Proven track record managing vendor relationships and complex third-party technology partnerships

  • Experience leading cross-functional teams and working across multiple business units

  • Strong communication and relationship management abilities

  • Demonstrated success implementing industry best practices and driving technical transformation

  • Experience with AI and emerging technologies in operations and reliability engineering

Nice to Have:

  • Payments industry experience

What's in it for you?

We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.

  • A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable

  • Leaders who support your development through coaching and managing opportunities

  • Ability to make a difference and lasting impact

  • Work in a dynamic, collaborative, progressive, and high-performing team

  • A world-class training program in financial services

  • Opportunities to do challenging work

  • Opportunities to take on progressively greater accountabilities

  • Opportunities to build close relationships with clients and stakeholders

#LI-POST

#TECHPJ

Job Skills

Application Development, Application Maintenance, Applications Architecture, Commercial Acumen, Enterprise Application Delivery, Information Technology Management, Information Technology Trends, Programming Languages, System Applications

Additional Job Details

Address:

RBC WATERPARK PLACE, 88 QUEENS QUAY W:TORONTO

City:

Toronto

Country:

Canada

Work hours/week:

37.5

Employment Type:

Full time

Platform:

TECHNOLOGY AND OPERATIONS

Job Type:

Regular

Pay Type:

Salaried

Posted Date:

2026-08-11

Application Deadline:

2026-09-14

Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above

Our Employment Opportunities

At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.

Join our Talent Community
Stay in-the-know about great career opportunities at RBC. Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you.
Expand your limits and create a new future together at RBC. Find out how we use our passion and drive to enhance the well-being of our clients and communities at jobs.rbc.com.

RBC is presently inviting candidates to apply for this existing vacancy. Applying to this posting allows you to express your interest in this current career opportunity at RBC. Qualified applicants may be contacted to review their resume in more detail.

Employment Type: FULL_TIME