1

Incident Problem Manager Jobs in Ontario (NOW HIRING)

Own the integrated ITSM audit delivery portfolio, covering Incident, Problem, Change, CMDB, Release, and ITAM findings. * Manage issues from Management Action Plan (MAP) creation through evidence ...

Participate in day-to-day operations including incident, problem, and change management * Support incident response and assist in root cause analysis and post-incident reviews * Maintain and follow ...

(CAN) Solution Consultant III

Mississauga, ON · On-site

CA$91K - CA$126K/yr

Support incident and problem management for existing technology in operations, working closely with operational teams to drive issue resolution and continuous improvement. * Manage and mentor ...

Demonstrated proficiency in Incident, Problem, and Change Management (INC/PRB/CHG) - understanding ... how these disciplines interconnect and how to leverage them to drive resiliency outcomes. Knowledge ...

Participate in incident, problem, and change management processes * Support incident triage, investigation, and documentation of root causes * Maintain and update runbooks, knowledge articles, and ...

Drive continuous improvement in incident, problem, and change management practices. * Manage vendor and managed service provider relationships, including contract performance and service delivery.

Collaborates effectively with engineering, operations, security, and vendor teams to support high-impact issues, follows established incident, problem, and change management processes, shares ...

Showing results 21-40

Incident Problem Manager information

How does an Incident Problem Manager typically collaborate with technical and non-technical teams during major incidents?

Incident Problem Managers play a critical role in bridging communication between technical teams (like IT support, network engineers, or developers) and non-technical stakeholders (such as business unit leaders or customer service). During major incidents, they coordinate response efforts, facilitate status updates, and ensure all parties are aligned on next steps and remediation plans. Effective collaboration involves translating complex technical issues into clear, actionable information for non-technical audiences, managing expectations, and driving post-incident reviews to prevent recurrence. This cross-functional coordination is essential for minimizing business impact and ensuring swift resolution.

What is the difference between Incident Problem Manager vs Incident Coordinator?

AspectIncident Problem ManagerIncident Coordinator
Primary RoleManages the lifecycle of incidents and problems to minimize impact and prevent recurrenceCoordinates incident response activities, ensuring timely resolution and communication
CertificationsITIL Foundation, Problem Management certificationsITIL Foundation, Incident Management certifications
Work EnvironmentTypically in IT service management teams, focusing on problem analysisOperational teams, focusing on incident handling and communication

While both roles are involved in incident management, the Incident Problem Manager focuses on identifying root causes and preventing future issues, whereas the Incident Coordinator handles day-to-day incident response and communication. Both roles are essential for effective IT service delivery but differ in scope and responsibilities.

What does an Incident Problem Manager do?

An Incident Problem Manager is responsible for overseeing the process of identifying, investigating, and resolving incidents and underlying problems within an organization's IT systems. They work to minimize the impact of disruptions, coordinate responses to incidents, and analyze root causes to prevent future occurrences. This role often involves collaborating with technical teams, managing communication with stakeholders, and ensuring that procedures are followed according to IT service management frameworks like ITIL. Their goal is to improve IT service reliability and reduce downtime for the business.

What are the key skills and qualifications needed to thrive as an Incident Problem Manager, and why are they important?

To thrive as an Incident Problem Manager, you need strong analytical skills, IT service management knowledge, and experience with incident and problem resolution processes, often supported by ITIL certification. Familiarity with ITSM tools like ServiceNow, Jira Service Management, or BMC Remedy is typically required. Exceptional communication, leadership, and critical thinking abilities enable effective coordination and root-cause analysis across teams. These skills are crucial to minimizing downtime, ensuring service continuity, and driving long-term improvements in IT operations.
What job categories do people searching Incident Problem Manager jobs in Ontario look for? The top searched job categories for Incident Problem Manager jobs in Ontario are:
What cities in Ontario are hiring for Incident Problem Manager jobs? Cities in Ontario with the most Incident Problem Manager job openings:
Infographic showing various Incident Problem Manager job openings in Ontario as of August 2026, with employment types broken down into 1% As Needed, 77% Full Time, 21% Part Time, and 1% Contract. Highlights an 93% Physical, 3% Hybrid, and 4% Remote job distribution.

Full-time

Re-posted 4 days ago


Job description

Purpose of Job
The Senior AI Platform Operations Engineer is accountable for the reliability, operability, and controlled enablement of the organization's AI platform. 

This role ensures that AI Platform services and solutions are production-ready, secure, observable, and compliant by executing disciplined operational practices, across platform management, monitoring, incident coordination, and governance control enforcement. 

The incumbent plays a key role in enabling the safe and scalable adoption of AI by ensuring that AI solutions are deployed, monitored, supported, and continuously improved in line with enterprise standards for reliability, security, and compliance

Main Activities

AI Platform Reliability and Operations
   Administer and operate the AI platform to ensure availability, performance, and resilience across environments, integrations, and supporting infrastructure.
   Monitor platform health using dashboards, logs, metrics, and alerts, and coordinate incident and service restoration activities.
   Lead operational triage, escalation coordination, and post-incident reviews to strengthen services stability and resilience. 
   Track and report on service reliability indicators, incident trends, and operational performance.

AI Platform Enablement & Production Readiness
   Enable approved AI use cases into production by ensuring: 
o    Environment readiness, 
o    Dependency validation, 
o    Completion of operational readiness checklists, 
o    Structured service transition activities
   Support platform lifecycle management through:
o    Release coordination, 
o    Change readiness validation, 
o    Maintenance and capacity planning.
   Ensure AI platform changes meet defined operational and control readiness criteria prior to release

Observability, Automation & AI Ops
   Implement and maintain observability capabilities, including telemetry, logging, metrics, and traces required for enterprise AI operations.
   Analyze operational data to identify anomalies, recurring issues, root-cause patterns.
   Implement AI Ops use cases such as:
o    Alert correlation, 
o    Anomaly detection, 
o    Root-cause support, 
o    Forecasting and predictive insights, 
o    Automation of repetitive operational tasks.
   Continuously improve operational efficiency through targeted automation and process optimization.

Governance, Risk, & Control Execution
   Execute governance controls for AI solutions, including:
o    Usage and access controls, 
o    Data privacy considerations, 
o    Auditability and traceability, 
o    Human oversight requirements
   Ensure operational practices align with enterprise security policies, risk controls, and compliance requirements.
   Maintain documentation and evidence required for audit, governance reviews, production readiness checkpoints, and control validation.
   Identify control gaps and escalate risks appropriately to relevant governance and risk stakeholders.

AI Asset Visibility & Operational Integrity
   Maintain operational visibility of AI platform assets required for monitoring, support, and cost alignment. 
   Validate asset ownership, relationships, and lifecycle status in collaboration with application and platform owners. 
   Support ongoing audits to ensure AI assets and associated cost attribution remain accurate and current.

Knowledge/Skill Requirements
   University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
   5-7 years of experience in platform operations, site reliability engineering, DevOps, cloud operations, or enterprise IT operations.
   Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting.
 

Technical Expertise:
   Experience with cloud platforms, observability, automation, configuration management, and integration patterns, including Azure Automation runbooks (PowerShell/Python), Azure AI, Copilot integrations, AKS, virtual networks (hub-and-spoke), and App Service.
   Expertise with observability tools such as Azure Monitor, Application Insights, and Grafana.
   Experience with CI/CD and automation tools such as Azure DevOps, GitHub Actions, and Logic Apps.
   Knowledge of configuration management and infrastructure-as-code tools such as Bicep, Terraform, Azure Policy, Key Vault, and relevant open-source technologies.
   Knowledge of integration and event-driven technologies such as API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka.
   Working knowledge of platform-supporting data and search services such as Elastic, Azure AI Search, and Cosmos DB.
   Knowledge of enterprise network, edge security, and related internal platforms such as DNA, Fortinet, and Akamai is an asset.

Additional Capabilities:
   Working knowledge of AI/ML operational concepts, including model lifecycle support, telemetry, governance controls, human-in-the-loop practices, and production monitoring.
   Strong understanding of ITIL/ITSM processes, including change, release, incident, problem, configuration, and service reporting practices.
   Analytical and structured thinker with strong troubleshooting, root-cause analysis, prioritization, and continuous improvement skills.
   Strong service orientation, professional maturity, and the ability to collaborate effectively across operations, engineering, security, risk, data, and business teams.
   Experience creating technical documentation, operational procedures, support playbooks, dashboards, and user guidance materials.
   Knowledge of security, privacy, audit, and compliance considerations relevant to enterprise AI and platform operations.
Job Complexities / Thinking Challenges
This role requires balancing platform reliability, operational efficiency, and governance discipline in a rapidly evolving AI environment. 

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
apply for this job