This role is responsible for designing, implementing, and maintaining SRE (Site Reliability Engineering) and AIOps (Artificial Intelligence for IT Operations) capabilities to ensure system ...
This role is responsible for designing, implementing, and maintaining SRE (Site Reliability Engineering) and AIOps (Artificial Intelligence for IT Operations) capabilities to ensure system ...
Senior Site Reliability Engineer
Toronto, ON · On-site
Drive operational excellence by championing SRE practices, automating repetitive workflows, and aligning technical initiatives with business priorities to improve system resilience * Provide ...
Senior Site Reliability Engineer
Toronto, ON · On-site
Drive operational excellence by championing SRE practices, automating repetitive workflows, and aligning technical initiatives with business priorities to improve system resilience * Provide ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON · Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service during incidents, and minimise customer impact. When issues occur, the role focuses not only on ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON · Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service during incidents, and minimise customer impact. When issues occur, the role focuses not only on ...
... (SRE) methodologies and technologies. Participate in defining standard reliability and resilience requirements for infrastructure and application components. Act as a point of escalation for ...
New
... (SRE) methodologies and technologies. Participate in defining standard reliability and resilience requirements for infrastructure and application components. Act as a point of escalation for ...
New
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON · Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
Quick apply
Reliability Expert - Fully Remote | Upto $120/hr
Toronto, ON · Remote
CA$120/hr
Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80-$120/hour Location: Remote Role Responsibilities * Evaluate AI-generated artifacts against domain-specific quality ...
... (SRE) methodologies and technologies. Participate in defining standard reliability and resilience requirements for infrastructure and application components. Act as a point of escalation for ...
New
... (SRE) methodologies and technologies. Participate in defining standard reliability and resilience requirements for infrastructure and application components. Act as a point of escalation for ...
New
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly ...
Head of Platform Engineering, Reliability & Control
Toronto, ON · Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
Head of Platform Engineering, Reliability & Control
Toronto, ON · Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and ...
The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and ...
The SRE will be responsible for aligning client requirements to standard features as well as deploying those features using automated technologies. Our SREs are pivotal in delivering technology ...
The SRE will be responsible for aligning client requirements to standard features as well as deploying those features using automated technologies. Our SREs are pivotal in delivering technology ...
Senior Site Reliability Engineer
Toronto, ON · On-site
We're searching for a Senior Site Reliability Engineer who not only brings deep technical expertise but also thrives on engaging with customers, partners, and cross-functional teams. In this customer ...
Senior Site Reliability Engineer
Toronto, ON · On-site
We're searching for a Senior Site Reliability Engineer who not only brings deep technical expertise but also thrives on engaging with customers, partners, and cross-functional teams. In this customer ...
Site Reliability Engineer - Interface & Connectivity
Toronto, ON · Hybrid
CA$68K - CA$103K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Site Reliability Engineer, you will join one of our Interface and Connectivity Product Areas and become part of a collaborative team focused on operating and ...
Site Reliability Engineer - Interface & Connectivity
Toronto, ON · Hybrid
CA$68K - CA$103K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Site Reliability Engineer, you will join one of our Interface and Connectivity Product Areas and become part of a collaborative team focused on operating and ...
Join RBC as the ATM Lead Site Reliability Engineer and lead the reliability, scalability, and performance of our critical ATM infrastructure. This hybrid role combines strategic leadership with hands ...
Join RBC as the ATM Lead Site Reliability Engineer and lead the reliability, scalability, and performance of our critical ATM infrastructure. This hybrid role combines strategic leadership with hands ...
Senior Site Reliability Engineer - Azure
Toronto, ON · Hybrid
CA$99K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Senior Site Reliability Engineer, you will be embedded within one of our Product Areas, taking ownership of specific responsibility domains where your experience ...
Senior Site Reliability Engineer - Azure
Toronto, ON · Hybrid
CA$99K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Senior Site Reliability Engineer, you will be embedded within one of our Product Areas, taking ownership of specific responsibility domains where your experience ...
Minimum of 5 years of experience in automating infrastructure, service delivery, and engineering site reliability, maintaining infrastructure on premise and in cloud environment * Product does not ...
Quick apply
Minimum of 5 years of experience in automating infrastructure, service delivery, and engineering site reliability, maintaining infrastructure on premise and in cloud environment * Product does not ...
Expert SRE
Mississauga, ON · Hybrid
We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability of our containerized business applications. You will bridge the gap between development and ...
Expert SRE
Mississauga, ON · Hybrid
We are seeking a Site Reliability Engineer to ensure the availability, performance, and reliability of our containerized business applications. You will bridge the gap between development and ...
Sr. Site Reliability Engineering - Azure Platform
Toronto, ON · Hybrid
CA$99K - CA$149K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Senior Site Reliability Engineer, you will be embedded within one of our Azure Product Areas. You will work closely with engineers, clients, and stakeholders to ...
Sr. Site Reliability Engineering - Azure Platform
Toronto, ON · Hybrid
CA$99K - CA$149K/yr
WHY THIS ROLE IS IMPORTANT TO US As a Senior Site Reliability Engineer, you will be embedded within one of our Azure Product Areas. You will work closely with engineers, clients, and stakeholders to ...
Establish and mature SRE practices across critical applications and platforms, including SLIs, SLOs, SLAs, service health indicators, post-incident reviews and continuous reliability improvement.
New
Establish and mature SRE practices across critical applications and platforms, including SLIs, SLOs, SLAs, service health indicators, post-incident reviews and continuous reliability improvement.
New
Prod Support - SRE
Toronto, ON · On-site
Catch point, Dynatrace, Kubernetes, Linux, Python, Red Hat, Open Shift, SRE We are looking for a Production support Engineer to work for the Commercial Line of Business. An ideal candidate should be ...
Prod Support - SRE
Toronto, ON · On-site
Catch point, Dynatrace, Kubernetes, Linux, Python, Red Hat, Open Shift, SRE We are looking for a Production support Engineer to work for the Commercial Line of Business. An ideal candidate should be ...
Role Overview: We're looking for a Site Reliability Engineer to improve the reliability, resilience, and operational readiness of our services. You'll work closely with engineering teams to improve ...
Role Overview: We're looking for a Site Reliability Engineer to improve the reliability, resilience, and operational readiness of our services. You'll work closely with engineering teams to improve ...
Associate Site Reliability Engineer information
What is the difference between Associate Site Reliability Engineer vs Software Engineer?
| Aspect | Associate Site Reliability Engineer | Software Engineer |
|---|---|---|
| Required Credentials | Bachelor's in CS or related field, some certifications (e.g., Linux, cloud) | Bachelor's in CS or related field, coding certifications often preferred |
| Work Environment | Operations-focused, maintaining system reliability, monitoring | Development-focused, designing and coding software applications |
| Employer & Industry Usage | Tech companies, cloud providers, SaaS firms | Tech companies, startups, software firms |
| Common Search & Comparison Intent | Understanding roles in reliability and operations | Understanding software development roles |
Associate Site Reliability Engineers focus on maintaining system stability, monitoring, and incident response, often working closely with operations teams. Software Engineers primarily develop, test, and deploy software applications. While both roles require strong technical skills and a background in computer science, their core responsibilities differ—one emphasizes reliability and system health, the other software creation.

Job description
Job Description
WHAT IS THE OPPORTUNITY?
This role is responsible for designing, implementing, and maintaining SRE (Site Reliability Engineering) and AIOps (Artificial Intelligence for IT Operations) capabilities to ensure system reliability, proactive monitoring, and automation of self-healing operations. In addition to day-to-day support, the position provides end-to-end operational ownership across systems managed by multiple enterprise teams, including incident coordination, dependency management, and escalation. The role is also responsible for key security and compliance functions such as service ID and certificate management, SSO updates, vulnerability remediation, and lifecycle management of end-of-life components.
Our team supports a portfolio of multi-platform HR data pipelines that move and process data into Snowflake through multiple integrated components, requiring end-to-end monitoring, coordination, and support across systems.
In addition, we support SaaS-based applications that are primarily vendor-managed, while we retain responsibility for integration, access management, monitoring, and operational oversight.
WHAT WILL YOU DO?
Key Responsibilities:
SRE & Reliability Engineering
- Define and operationalize SLIs, SLOs, and error budgets
- Own the incident management lifecycle (detection triage resolution RCA prevention)
- Lead problem management and eliminate recurring issues
- Develop and maintain runbooks, playbooks, and recovery procedures
- Drive resilience engineering, including failover testing and capacity planning
AIOps, Observability & Logging
- Implement and optimize AIOps capabilities using platforms such as Moogsoft
- Leverage Dynatrace for deep APM insights
- Integrate alerting and escalation workflows with PagerDuty
- Utilize synthetic monitoring via Catchpoint
- Design and maintain centralized logging solutions using the ELK stack (Elasticsearch, Logstash, Kibana) and enterprise Logging as a Service (LaaS) platforms
- Perform log analysis, correlation, and anomaly detection to support proactive issue identification
- Drive event noise reduction and intelligent alerting strategies
- Build dashboards and observability KPIs for operational insights
Automation & Self-Healing Systems
- Design and implement automation-first solutions using Ansible and scripting (Python, Bash)
- Enable self-healing capabilities (auto-remediation, restart logic, workflow recovery)
- Orchestrate workflows using Stonebranch
- Reduce operational toil through automation and continuous improvement
Application & Data Platform Support
- Provide L2/L3 support for data pipelines and integration workflows, including Snowflake ingestion and transformation processes
- Support Snowflake pipelines, ETL workflows, and orchestration dependencies
- Troubleshoot across distributed systems including:
- Object storage (e.g., S3)
- APIs, messaging, and file transfer systems
- Support containerized workloads on OpenShift
Infrastructure & Platform Expertise
- Administer and support Windows Server and IIS-based applications
- Manage certificate lifecycle (TLS, SAML, OAuth)
- Support SSO integrations via Microsoft Entra ID
- Work with relational databases (SQL Server, PostgreSQL) for troubleshooting and performance tuning
Disaster Recovery & Resiliency
- Lead and coordinate Disaster Recovery (DR) planning and execution
- Validate failover processes across dependent systems
- Ensure end-to-end DR readiness across integrated platforms
- Document recovery strategies and participate in DR exercises
Governance, Compliance & Collaboration
- Ensure adherence to enterprise security, compliance, and audit requirements
- Collaborate with IAM, PAM, logging, and cloud platform teams
- Support change management, CAB processes, and release coordination
- Provide reporting, KPIs, and executive summaries on system health
WHAT DO YOU NEED TO SUCCEED?
Must Have:
- 3+ years of SRE or Systems Engineering experience with strong technical expertise.
- Experience with ServiceNow, ITSM processes including incident, problem, change, and release management.
- Demonstrated ability to work independently, take ownership, and drive projects to completion.
- Knowledge of containerized platforms such as Kubernetes or OpenShift.
- Experience working with Ansible Automation Platform or strong willingness to learn.
- In-depth knowledge of monitoring and observability tools such as Dynatrace, Elasticsearch, Moogsoft, Catchpoint, and PagerDuty.
- Knowledge and experience with scripting languages such as BASH, Python and PowerShell.
- Experience with Linux and Windows Server administration.
- Experience with or strong interest in intelligent monitoring, anomaly detection, and automation technologies
- Excellent problem-solving skills and attention to detail.
- Experience with API's and secure file transfer procedures.
Nice to Have:
- Experience with telemetry standardization (OpenTelemetry) and observability data correlation.
- Understanding of AI/ML concepts and their application to observability and operations (AIOps).
- Experience with tools such as Ansible, Stonebranch, Kafka, and their role in system reliability.
- Experience with CI/CD and developer platform tools such as Jenkins, GitHub, GitHub Actions, Artifactory, and Vault.
- Familiarity with Snowflake and MongoDB Atlas, including experience with writing and executing basic queries, validating data, and troubleshooting production issues.
- Understanding of Single Sign-On(SSO) technologies, including Microsoft Entra ID, SAML and OAuth.
- Experience supporting enterprise applications in banking or financial services industry with understanding of regulatory, security and compliance requirements.
WHAT'S IN IT FOR YOU?
We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.
A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
Leaders who support your development through coaching and managing opportunities
Ability to make a difference and lasting impact
Work in a dynamic, collaborative, progressive, and high-performing team
A world-class training program in financial services
Opportunities to do challenging work
#LI-POST
#TECHPJ
Job Skills
Critical Thinking, Customer Support Systems, Group Problem Solving, Installation Support, IT Service Level Management, IT Service Management (ITSM), IT Standards, Technical TroubleshootingAdditional Job Details
Address:
City:
Country:
Work hours/week:
Employment Type:
Platform:
Job Type:
Pay Type:
Posted Date:
Application Deadline:
Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above
Our Employment Opportunities
At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.
Join our Talent Community
Stay in-the-know about great career opportunities at RBC. Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you.
Expand your limits and create a new future together at RBC. Find out how we use our passion and drive to enhance the well-being of our clients and communities at jobs.rbc.com.
RBC is presently inviting candidates to apply for this existing vacancy. Applying to this posting allows you to express your interest in this current career opportunity at RBC. Qualified applicants may be contacted to review their resume in more detail.
Employment Type: FULL_TIMEAbout Royal Bank of Canada
Sourced by ZipRecruiter
Industry
Banking and credit intermediation
Company size
10,000+ Employees
Headquarters location
Toronto, Ontario, CA