We are seeking a highly skilled and experienced Observability Engineer to lead and evolve our monitoring and observability capabilities across complex, hybrid environments. This role requires deep ...
We are seeking a highly skilled and experienced Observability Engineer to lead and evolve our monitoring and observability capabilities across complex, hybrid environments. This role requires deep ...
By measuring and monitoring our operations you find opportunities to improve our systems in order ... Collaborate with cross-functional engineering teams to enhance Lyft's observability and meet ...
By measuring and monitoring our operations you find opportunities to improve our systems in order ... Collaborate with cross-functional engineering teams to enhance Lyft's observability and meet ...
Senior / Staff Software Engineer (Observability / SRE)
Toronto, ON · On-site +1
CA$148K - CA$249K/yr
... observability stack, used to ... monitor the health and performance of cloud and on-prem environments. - Develop and extend ...
Senior / Staff Software Engineer (Observability / SRE)
Toronto, ON · On-site +1
CA$148K - CA$249K/yr
... observability stack, used to ... monitor the health and performance of cloud and on-prem environments. - Develop and extend ...
Senior Instrumentation and Monitoring Engineer
Markham, ON · On-site
CA$104K - CA$138K/yr
Review, analyze, and interpret monitoring data to support informed engineering decisions. * Set up, maintain, and manage monitoring data visualization platforms. * Prepare monitoring reports, cost ...
Senior Instrumentation and Monitoring Engineer
Markham, ON · On-site
CA$104K - CA$138K/yr
Review, analyze, and interpret monitoring data to support informed engineering decisions. * Set up, maintain, and manage monitoring data visualization platforms. * Prepare monitoring reports, cost ...
Senior Instrumentation and Monitoring Engineer
Markham, ON · On-site
CA$104K - CA$138K/yr
Review, analyze, and interpret monitoring data to support informed engineering decisions. * Set up, maintain, and manage monitoring data visualization platforms. * Prepare monitoring reports, cost ...
Senior Instrumentation and Monitoring Engineer
Markham, ON · On-site
CA$104K - CA$138K/yr
Review, analyze, and interpret monitoring data to support informed engineering decisions. * Set up, maintain, and manage monitoring data visualization platforms. * Prepare monitoring reports, cost ...
Senior Software Engineer
Toronto, ON · On-site
Knowledge of monitoring concepts (logs, metrics, traces, telemetry) and observability tools ... Programming Language), React.js, Software Development Life Cycle (SDLC), SRE Observability ...
Senior Software Engineer
Toronto, ON · On-site
Knowledge of monitoring concepts (logs, metrics, traces, telemetry) and observability tools ... Programming Language), React.js, Software Development Life Cycle (SDLC), SRE Observability ...
Head of Platform Engineering, Reliability & Control
Toronto, ON · Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
Head of Platform Engineering, Reliability & Control
Toronto, ON · Hybrid
CA$170K - CA$185K/yr
Define and implement SRE and observability standards, enabling proactive monitoring, faster incident detection, improved root cause analysis, and consistent reliability practices across distributed ...
Senior Platform Engineer
Vaughan, ON · On-site
Enable Observability: Ensure all infrastructure and pipelines emit the right telemetry for downstream monitoring by our SRE/Observability team. * Cross-Functional Partnership: Partner closely with ...
Senior Platform Engineer
Vaughan, ON · On-site
Enable Observability: Ensure all infrastructure and pipelines emit the right telemetry for downstream monitoring by our SRE/Observability team. * Cross-Functional Partnership: Partner closely with ...
Lead Data Engineer
Toronto, ON · On-site
Ingestion, orchestration, and observability * Engineer resilient, observable ingestion patterns intoS3/Snowflake/RDS. * Orchestrate pipelines using Airflow (or equivalent) and/or AWS-native services ...
Lead Data Engineer
Toronto, ON · On-site
Ingestion, orchestration, and observability * Engineer resilient, observable ingestion patterns intoS3/Snowflake/RDS. * Orchestrate pipelines using Airflow (or equivalent) and/or AWS-native services ...
... Monitoring (APM), and Site Reliability Engineering (SRE)-ensuring their priorities, investments ... Platform Optimization, Cost & Scale - Optimize observability platform efficiency, performance, and ...
... Monitoring (APM), and Site Reliability Engineering (SRE)-ensuring their priorities, investments ... Platform Optimization, Cost & Scale - Optimize observability platform efficiency, performance, and ...
AI ModelOps Engineer
Toronto, ON · On-site
Design and implement observability, monitoring, alerting, and troubleshooting capabilities to ... Collaborate with data scientists, AI engineers, cloud engineers, architects, software developers ...
New
AI ModelOps Engineer
Toronto, ON · On-site
Design and implement observability, monitoring, alerting, and troubleshooting capabilities to ... Collaborate with data scientists, AI engineers, cloud engineers, architects, software developers ...
New
AI ModelOps Engineer
Toronto, ON · On-site
Design and implement observability, monitoring, alerting, and troubleshooting capabilities to ... Collaborate with data scientists, AI engineers, cloud engineers, architects, software developers ...
New
AI ModelOps Engineer
Toronto, ON · On-site
Design and implement observability, monitoring, alerting, and troubleshooting capabilities to ... Collaborate with data scientists, AI engineers, cloud engineers, architects, software developers ...
New
AI ModelOps Engineer
Toronto, ON · On-site
As an AI ModelOps Engineer, you will be part of the Enterprise AI Platforms & AI ModelOps team ... Design and implement observability, monitoring, alerting, and troubleshooting capabilities to ...
AI ModelOps Engineer
Toronto, ON · On-site
As an AI ModelOps Engineer, you will be part of the Enterprise AI Platforms & AI ModelOps team ... Design and implement observability, monitoring, alerting, and troubleshooting capabilities to ...
Platform Engineer Position Overview We are seeking a Platform Engineer to support the development ... Observability & Support: Configure monitoring, logging, dashboards, and alerts. Troubleshoot ...
Platform Engineer Position Overview We are seeking a Platform Engineer to support the development ... Observability & Support: Configure monitoring, logging, dashboards, and alerts. Troubleshoot ...
Director, Global Fraud Technology - Service Reliability & Production Engineering
Toronto, ON · On-site
... ) principles across fraud technology platforms. * Establish reliability engineering practices, operational metrics, and service-level objectives. * Drive implementation of monitoring, observability ...
Director, Global Fraud Technology - Service Reliability & Production Engineering
Toronto, ON · On-site
... ) principles across fraud technology platforms. * Establish reliability engineering practices, operational metrics, and service-level objectives. * Drive implementation of monitoring, observability ...
As Director, DevOps, you will set the strategy, architecture, roadmap and operating model for CPP ... Maintain accountability for platform availability, resiliency, observability, capacity, incident ...
As Director, DevOps, you will set the strategy, architecture, roadmap and operating model for CPP ... Maintain accountability for platform availability, resiliency, observability, capacity, incident ...
Platform Engineering Manager
Toronto, ON · Hybrid
CA$140K - CA$160K/yr
This role will help build common patterns for observability, orchestration, data/control workflows ... Experience with Azure, Snowflake, Databricks, cloud-native monitoring, operational data platforms ...
Platform Engineering Manager
Toronto, ON · Hybrid
CA$140K - CA$160K/yr
This role will help build common patterns for observability, orchestration, data/control workflows ... Experience with Azure, Snowflake, Databricks, cloud-native monitoring, operational data platforms ...
... platform observability and capacity management. * Support scaling and monitoring of AI ... Experience with DevOps/Platform Engineering/SRE principles and designing for operational excellence.
... platform observability and capacity management. * Support scaling and monitoring of AI ... Experience with DevOps/Platform Engineering/SRE principles and designing for operational excellence.
Infrastructure Engineer
Toronto, ON · Remote
CA$140K - CA$190K/yr
... observability, secrets management, networking, data infrastructure, and deployment automation ... Monitor infrastructure utilization and cost, identifying opportunities to improve efficiency ...
Infrastructure Engineer
Toronto, ON · Remote
CA$140K - CA$190K/yr
... observability, secrets management, networking, data infrastructure, and deployment automation ... Monitor infrastructure utilization and cost, identifying opportunities to improve efficiency ...
Observability & Reliability * Experience designing and operating monitoring, logging, tracing, and ... Software Engineering & Platform Development * Strong proficiency in TypeScript and Node.js for ...
Observability & Reliability * Experience designing and operating monitoring, logging, tracing, and ... Software Engineering & Platform Development * Strong proficiency in TypeScript and Node.js for ...
Expert Observability Engineer
Mississauga, ON • Hybrid
Full-time
Medical, Life, Retirement, PTO
Re-posted 13 days ago
Job description
At Finastra, we're a global leader in financial services software, dedicated to expanding access to financial services and shaping what's next for the industry. Our technology powers missioncritical solutions across Lending, Payments and Universal Banking, supporting over 7,000 customers, including 80% of the world's top 50 banks, in more than 110 countries.
We are seeking a highly skilled and experienced Observability Engineer to lead and evolve our monitoring and observability capabilities across complex, hybrid environments. This role requires deep technical expertise in tools like Grafana, Site24x7, Azure Monitor, and a strong understanding of Azure Cloud, Windows, and Linux operating systems. The ideal candidate will also bring leadership experience, guiding teams in implementing scalable observability strategies that drive performance, reliability, and resilience.
Responsibilities & Deliverables
Design and implement observability solutions using Grafana, Site24x7, Azure Monitor, and other relevant tools.
Develop and maintain dashboards, alerts, and metrics to monitor infrastructure, applications, and services across cloud and on-prem environments.
Collaborate with cross-functional teams to define SLIs/SLOs and improve system reliability and performance.
Lead initiatives to enhance telemetry, logging, and tracing across distributed systems.
Provide technical leadership and mentorship to junior engineers and cross-functional teams.
Troubleshoot and resolve complex issues in real-time, leveraging deep knowledge of OS internals (Windows/Linux) and cloud infrastructure.
Drive adoption of best practices in monitoring, incident response, and post-mortem analysis.
Stay current with industry trends and emerging technologies in observability and resilience engineering.
Required Qualifications
Proficiency in observability, monitoring, or site reliability engineering.
Expert-level proficiency in Grafana (including custom dashboards, integrations, and alerting).
Hands-on experience with Site24x7, Azure Monitor, and other observability platforms.
Strong understanding of Azure Cloud architecture, services, and deployment models.
Deep technical knowledge of Windows and Linux operating systems, including performance tuning and diagnostics.
Experience with scripting and automation (e.g., PowerShell, Bash, Python).
Familiarity with containerized environments (Docker, Kubernetes) is a plus.
Excellent communication and team leadership skills, with a track record of mentoring and guiding technical teams.
Preferred Qualifications
Certifications in Azure or related technologies.
Experience with OpenTelemetry, Prometheus, ELK stack, or similar tools.
Background in resilience engineering or incident management frameworks.
As part of our hiring process, we may use artificial intelligence (AI) technology to help screen and shortlist applications. All final hiring decisions are made by our recruitment team.
Compensation: 112 - 130k
We are proud to offer a range of incentives to our employees worldwide. These benefits are available to everyone, regardless of grade, and reflect the values we stand for:
Flexibility: Enjoy unlimited vacation, subject to local regulations and business priorities. Benefit from hybrid working arrangements and inclusive policies such as paid time off for voting, bereavement, and sick leave.
Wellbeing: Access confidential onetoone support through our Employee Assistance Program, connect with our network of Wellbeing Champions and Gather Groups, and take part in monthly events and initiatives designed to help you thrive-inside and outside of work.
Health & Financial Security: Medical, life and disability insurance, retirement plans, lifestyle, and other benefits.*
Sustainability: Paid time off for volunteering and donationmatching opportunities to support causes that matter to you.
Inclusion: Get involved in our inclusion communities, such as Count Me In, Culture@Finastra, Proud@Finastra, Disabilities@Finastra, and Women@Finastra-open to everyone who wants to participate and contribute.
Career Development: Access online learning and accredited courses through our Skills & Career Navigator tool.
Recognition: Take part in our global recognition program, Finastra Celebrates, and share your voice through regular employee surveys that help shape our culture and ways of working.
*Specific benefits may vary by location.
At Finastra, each individual is unique-bringing their own ideas, perspectives, cultural backgrounds, and experiences. We learn from one another, value what makes us different, and create an environment where everyone feels included, supported, and able to be their authentic selves.
Be unique. Be exceptional. Help us make a difference at Finastra.
Finastra is committed to providing accessible employment practices that are in compliance with the Accessibility for Ontarians with Disabilities Act (AODA). We will accommodate applicants' needs upon request, throughout all stages of the recruitment process. Please inform us of the accommodation(s) that you may require. Information received related to accommodation will be addressed confidentially.