Must-have * 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or ... Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or ...
Must-have * 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or ... Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or ...
Lead Site Reliability Administrator
Mississauga, ON · On-site
CA$177K/yr
As a Site Reliability Engineer (SRE) at OpenText, you will be responsible for ensuring the ... Azure DevOps certification or equivalent certifications. * Familiarity with observability tools ...
New
Lead Site Reliability Administrator
Mississauga, ON · On-site
CA$177K/yr
As a Site Reliability Engineer (SRE) at OpenText, you will be responsible for ensuring the ... Azure DevOps certification or equivalent certifications. * Familiarity with observability tools ...
New
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level ... and engineering bottlenecks. * Establish and continuously improve observability, monitoring ...
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level ... and engineering bottlenecks. * Establish and continuously improve observability, monitoring ...
Head of Platform Engineering, Reliability & Control
Toronto, ON · Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
Head of Platform Engineering, Reliability & Control
Toronto, ON · Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level ... engineering bottlenecks. Establish and continuously improve observability, monitoring, alerting ...
Quick apply
Define and mature SRE practices, including Service Level Objectives (SLOs), Service Level ... engineering bottlenecks. Establish and continuously improve observability, monitoring, alerting ...
The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability ... Implement and manage observability solutions (logging, metrics, tracing) using OpenTelemetry ...
The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability ... Implement and manage observability solutions (logging, metrics, tracing) using OpenTelemetry ...
Join our Credit Technology team as a Lead Site Reliability Engineer, where you'll play a key role in drive operational excellence through technology, process optimization, and cross-functional ...
Join our Credit Technology team as a Lead Site Reliability Engineer, where you'll play a key role in drive operational excellence through technology, process optimization, and cross-functional ...
The Site Reliability Engineer designs and implements solutions to reduce toil and ensure reliability of our critical services at Sectigo. This is a full-time position, working in a hybrid model ...
Quick apply
The Site Reliability Engineer designs and implements solutions to reduce toil and ensure reliability of our critical services at Sectigo. This is a full-time position, working in a hybrid model ...
As a SRE, you will implement, measure and gather insights from Operational Level Indicators identifying areas for service improvements covering availability, performance, resilience, incidents and ...
As a SRE, you will implement, measure and gather insights from Operational Level Indicators identifying areas for service improvements covering availability, performance, resilience, incidents and ...
Senior Site Reliability Engineer
Kitchener, ON · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Quick apply
Senior Site Reliability Engineer
Kitchener, ON · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Senior Site Reliability Engineer
Ottawa, ON · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Quick apply
Senior Site Reliability Engineer
Ottawa, ON · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Senior Site Reliability Engineer
Kitchener, ON · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Quick apply
Senior Site Reliability Engineer
Kitchener, ON · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Anticipate, detect, and address production-related issues proactively using advanced monitoring, observability, and automation. Define, educate, and enforce SRE processes, guidelines, and practices ...
Anticipate, detect, and address production-related issues proactively using advanced monitoring, observability, and automation. Define, educate, and enforce SRE processes, guidelines, and practices ...
We are looking for a Site Reliability Engineer to help design and deploy, and operate the platforms ... Build and improve observability across monitoring, logging, and alerting * Partner closely with ...
We are looking for a Site Reliability Engineer to help design and deploy, and operate the platforms ... Build and improve observability across monitoring, logging, and alerting * Partner closely with ...
Senior Site Reliability Engineer
Ottawa, ON · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Quick apply
Senior Site Reliability Engineer
Ottawa, ON · Remote
$95K - $100K/yr
We are looking for an experienced Senior Site Reliability Engineer for our client. This is a permanent position that is remote to start with later relocation to Calgary or Winnipeg . Our client is a ...
Experience with enterprise monitoring and observability tools : Dynatrace Managed, Catchpoint ... Hands-on experience applying SRE principles (SLOs, error budgets, blameless postmortems)
Experience with enterprise monitoring and observability tools : Dynatrace Managed, Catchpoint ... Hands-on experience applying SRE principles (SLOs, error budgets, blameless postmortems)
Operate and maintain observability infrastructure including Grafana, alerts, dashboards, and log ... and SRE practices; * Managed production Kubernetes environments at scale: not just deployed ...
Operate and maintain observability infrastructure including Grafana, alerts, dashboards, and log ... and SRE practices; * Managed production Kubernetes environments at scale: not just deployed ...
Operate and maintain observability infrastructure including Grafana, alerts, dashboards, and log ... and SRE practices; * Managed production Kubernetes environments at scale: not just deployed ...
Operate and maintain observability infrastructure including Grafana, alerts, dashboards, and log ... and SRE practices; * Managed production Kubernetes environments at scale: not just deployed ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON · Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service ... Design, implement, and maintain monitoring and observability solutions, including dashboards ...
(Canada) - Junior Site Reliability Engineer
Mississauga, ON · Hybrid
CA$70K - CA$80K/yr
The Site Reliability Engineer works to ensure services operate smoothly, quickly restore service ... Design, implement, and maintain monitoring and observability solutions, including dashboards ...
Operate and maintain observability infrastructure including Grafana, alerts, dashboards, and log ... and SRE practices; * Managed production Kubernetes environments at scale: not just deployed ...
Operate and maintain observability infrastructure including Grafana, alerts, dashboards, and log ... and SRE practices; * Managed production Kubernetes environments at scale: not just deployed ...
Observability Site Reliability Engineer information
What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?
| Aspect | Observability Site Reliability Engineer | Monitoring Engineer |
|---|---|---|
| Focus | Ensuring system reliability through observability, automation, and incident response | Implementing and managing monitoring tools and dashboards |
| Skills | Cloud platforms, scripting, incident management, observability tools | Monitoring tools, alerting systems, data analysis |
| Work Environment | DevOps teams, cloud infrastructure, large-scale systems | Operations teams, infrastructure monitoring |
While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.
What are popular job titles related to Observability Site Reliability Engineer jobs in Ontario?
For Observability Site Reliability Engineer jobs in Ontario, the most frequently searched job titles are:
What job categories do people searching Observability Site Reliability Engineer jobs in Ontario look for?
The top searched job categories for Observability Site Reliability Engineer jobs in Ontario are:
What cities in Ontario are hiring for Observability Site Reliability Engineer jobs?
Cities in Ontario with the most Observability Site Reliability Engineer job openings:
Full-time
Posted 15 days ago
Job description
Job Description
What is the opportunity?
Join the Platform Engineering & AI Operations team within OTK0, where you'll sit at the intersection of Site Reliability Engineering and intelligent infrastructure operations. This role offers the chance to shape how the bank operates, monitors, and self-heals its private and public cloud platforms - from OpenShift clusters and Kafka environments to self-healing automation systems. You'll work on real problems at enterprise scale: reducing toil for NOC, Data Center, and Branch teams, building automation that eliminates manual work, and establishing reliable operational practices. If you want to move beyond traditional ops into the future of intelligent, autonomous infrastructure operations - this is the role.
What will you do?
- Support highly scalable, secure, and highly available architectures across private and public cloud platforms (Kubernetes/OpenShift, ECE, Confluent Kafka).
- Write code and scripts to automate infrastructure workflows and eliminate toil, including automation pipelines that reduce manual intervention across Data Center, Branch, and NOC operations.
- Extend self-healing automation capabilities built on Ansible, automating routine operational tasks (e.g., CPU remediation) to reduce manual intervention.
- Participate in and lead design reviews for new platform features, infrastructure changes, and operational integration points, ensuring alignment with security, reliability, and regulatory requirements.
- Collaborate with platform teams to provide technical feedback, contribute code changes to shared repositories, and establish data standards and pipelines (e.g., ServiceNow, Prometheus) that support operational excellence.
- Drive automation, CI/CD, and Infrastructure as Code practices across the team, leveraging Ansible and Terraform for deployment validation and self-healing remediation workflows.
- Minimize risk of reliability failures related to durability, availability, performance, and correctness, leveraging proactive alerting and anomaly detection.
- Participate in on-call rotation for platform support, incident management, and troubleshooting, triaging incidents via Grafana, Prometheus, Dynatrace, and PagerDuty.
What do you need to succeed?
Must-have
- 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or infrastructure operations.
- Strong working knowledge of Kubernetes/OpenShift administration and troubleshooting in enterprise environments.
- Hands-on experience with Ansible and Terraform for Infrastructure as Code and automation.
- Proficiency in Python scripting (core to infrastructure automation and platform development).
- Hands-on experience with monitoring and observability stacks (Prometheus, Grafana, ELK, or equivalent).
- Experience with incident management processes, on-call rotations, and post-incident review practices.
- Familiarity with capacity planning, threshold-based alerting, and performance trend analysis.
- Understanding of security and compliance fundamentals, including vulnerability assessment and remediation tracking.
- Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
Nice-to-have
- Experience with AI/ML concepts applied to operations (anomaly detection, intelligent alerting, predictive capacity planning).
- Hands-on experience with public cloud platforms (AWS, Azure, GCP) in hybrid or multi-cloud environments.
- Experience with GPU/compute infrastructure for ML inference workloads
What's in it for you?
We thrive on the challenge to be our best, progressive thinking to keep growing, and working together to deliver trusted advice to help our clients thrive and communities prosper. We care about each other, reaching our potential, making a difference to our communities, and achieving success that is mutual.
- A comprehensive Total Rewards Program including bonuses and flexible benefits, competitive compensation, commissions, and stock where applicable
- Leaders who support your development through coaching and managing opportunities
- Ability to make a difference and lasting impact
- Work in a dynamic, collaborative, progressive, and high-performing team
- Flexible work/life balance options
- Opportunities to do challenging work
- Opportunities to take on progressively greater accountabilities
- Access to a variety of job opportunities across business
#LI-post
#TECHPJ
Job Skills
Agile Methodology, Ansible Tower, Group Problem Solving, IT System Administration, IT Systems Integration, Kubernetes, Linux, Organizational Leadership, Product Services, RedHat OpenShift Administration, Red Hat OS Administration, Software Development Life Cycle (SDLC), System Applications, System Integration Testing (SIT), Systems SoftwareAdditional Job Details
Address:
City:
Country:
Work hours/week:
Employment Type:
Platform:
Job Type:
Pay Type:
Posted Date:
Application Deadline:
Note: Applications will be accepted until 11:59 PM on the day prior to the application deadline date above
Our Employment Opportunities
At RBC, we are guided by living shared values of Client First, Integrity, Collaboration, Respect and Excellence and winning together as One RBC. We believe an inclusive workplace that has diverse perspectives is core to our continued growth as one of the largest and most successful banks in the world. Maintaining a workplace where our employees feel supported to perform at their best, effectively collaborate, drive innovation, and grow professionally helps to bring our Purpose to life and create value for our clients and communities. RBC strives to deliver this through policies and programs intended to foster a workplace based on respect, belonging and opportunity for all.
Join our Talent Community
Stay in-the-know about great career opportunities at RBC. Sign up and get customized info on our latest jobs, career tips and Recruitment events that matter to you.
Expand your limits and create a new future together at RBC. Find out how we use our passion and drive to enhance the well-being of our clients and communities at jobs.rbc.com.
RBC is presently inviting candidates to apply for this existing vacancy. Applying to this posting allows you to express your interest in this current career opportunity at RBC. Qualified applicants may be contacted to review their resume in more detail.
Employment Type: FULL_TIMEAbout Royal Bank of Canada
Sourced by ZipRecruiter
Industry
Banking and credit intermediation
Company size
10,000+ Employees
Headquarters location
Toronto, Ontario, CA