Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting. Technical Expertise:
Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting. Technical Expertise:
Lead incident, problem and change management for database platforms, ensuring strict SLAs adherence. * Establish proactive monitoring, capacity planning, performance tuning and disaster recovery ...
Lead incident, problem and change management for database platforms, ensuring strict SLAs adherence. * Establish proactive monitoring, capacity planning, performance tuning and disaster recovery ...
Senior Manager, ITSM Audit
Toronto, ON · On-site
Own the integrated ITSM audit delivery portfolio, covering Incident, Problem, Change, CMDB, Release, and ITAM findings. * Manage issues from Management Action Plan (MAP) creation through evidence ...
Senior Manager, ITSM Audit
Toronto, ON · On-site
Own the integrated ITSM audit delivery portfolio, covering Incident, Problem, Change, CMDB, Release, and ITAM findings. * Manage issues from Management Action Plan (MAP) creation through evidence ...
Lead incident, problem and change management for database platforms, ensuring strict SLAs adherence. * Establish proactive monitoring, capacity planning, performance tuning and disaster recovery ...
Lead incident, problem and change management for database platforms, ensuring strict SLAs adherence. * Establish proactive monitoring, capacity planning, performance tuning and disaster recovery ...
Lead System Administrator
Toronto, ON · On-site
Incident and Problem Management including during and post-incident analysis. * Researching and implementing new technology per project and support needs. * Set-up and support of workflow and job ...
Lead System Administrator
Toronto, ON · On-site
Incident and Problem Management including during and post-incident analysis. * Researching and implementing new technology per project and support needs. * Set-up and support of workflow and job ...
Senior SRE/AIOps Engineer
Toronto, ON · On-site
Own the incident management lifecycle (detection triage resolution RCA prevention) * Lead problem management and eliminate recurring issues * Develop and maintain runbooks, playbooks, and recovery ...
Senior SRE/AIOps Engineer
Toronto, ON · On-site
Own the incident management lifecycle (detection triage resolution RCA prevention) * Lead problem management and eliminate recurring issues * Develop and maintain runbooks, playbooks, and recovery ...
Own the service-management discipline -- incident, problem, and change management, plus request fulfillment and onboarding/offboarding workflows. * Lead knowledge management -- author, edit, and QA ...
Quick apply
Own the service-management discipline -- incident, problem, and change management, plus request fulfillment and onboarding/offboarding workflows. * Lead knowledge management -- author, edit, and QA ...
Knowledge of ITIL philosophy would be an asset (Service Desk Management, Incident & Problem Management and Change Management) * Available to work varied shifts within a 24/7 environment Please note ...
Knowledge of ITIL philosophy would be an asset (Service Desk Management, Incident & Problem Management and Change Management) * Available to work varied shifts within a 24/7 environment Please note ...
Cloud Engineer
Toronto, ON · On-site +1
Participate in day-to-day operations including incident, problem, and change management * Support incident response and assist in root cause analysis and post-incident reviews * Maintain and follow ...
Cloud Engineer
Toronto, ON · On-site +1
Participate in day-to-day operations including incident, problem, and change management * Support incident response and assist in root cause analysis and post-incident reviews * Maintain and follow ...
(CAN) Solution Consultant III
Mississauga, ON · On-site
CA$91K - CA$126K/yr
Support incident and problem management for existing technology in operations, working closely with operational teams to drive issue resolution and continuous improvement. * Manage and mentor ...
(CAN) Solution Consultant III
Mississauga, ON · On-site
CA$91K - CA$126K/yr
Support incident and problem management for existing technology in operations, working closely with operational teams to drive issue resolution and continuous improvement. * Manage and mentor ...
E&P Resiliency Manager
Toronto, ON · Hybrid
Demonstrated proficiency in Incident, Problem, and Change Management (INC/PRB/CHG) - understanding ... how these disciplines interconnect and how to leverage them to drive resiliency outcomes. Knowledge ...
E&P Resiliency Manager
Toronto, ON · Hybrid
Demonstrated proficiency in Incident, Problem, and Change Management (INC/PRB/CHG) - understanding ... how these disciplines interconnect and how to leverage them to drive resiliency outcomes. Knowledge ...
This position will be hands-on, providing production application support services to the Fixed Income Trading Business, leveraging ITIL service management principles of Event, Incident, Problem and ...
This position will be hands-on, providing production application support services to the Fixed Income Trading Business, leveraging ITIL service management principles of Event, Incident, Problem and ...
Participate in incident, problem, and change management processes * Support incident triage, investigation, and documentation of root causes * Maintain and update runbooks, knowledge articles, and ...
Participate in incident, problem, and change management processes * Support incident triage, investigation, and documentation of root causes * Maintain and update runbooks, knowledge articles, and ...
Technical Lead - Application Development & Maintenance
Mississauga, ON · On-site +1
CA$76K - CA$179K/yr
Ensure SLA/KPI adherence for incident, change, and problem management * Design and implement integration flows using OSB and web services * Manage release cycles and deployments across environments
Technical Lead - Application Development & Maintenance
Mississauga, ON · On-site +1
CA$76K - CA$179K/yr
Ensure SLA/KPI adherence for incident, change, and problem management * Design and implement integration flows using OSB and web services * Manage release cycles and deployments across environments
Director - IT Operations
Gloucester, ON · On-site
Drive continuous improvement in incident, problem, and change management practices. * Manage vendor and managed service provider relationships, including contract performance and service delivery.
Quick apply
Director - IT Operations
Gloucester, ON · On-site
Drive continuous improvement in incident, problem, and change management practices. * Manage vendor and managed service provider relationships, including contract performance and service delivery.
Incident & Problem Management * Participate in major incident management, providing technical leadership and structured analysis. * Contribute to or lead root cause analysis (RCA) and post-incident ...
Incident & Problem Management * Participate in major incident management, providing technical leadership and structured analysis. * Contribute to or lead root cause analysis (RCA) and post-incident ...
Incident & Problem Management * Participate in major incident management, providing technical leadership and structured analysis. * Contribute to or lead root cause analysis (RCA) and post-incident ...
Quick apply
Incident & Problem Management * Participate in major incident management, providing technical leadership and structured analysis. * Contribute to or lead root cause analysis (RCA) and post-incident ...
... Incident Management, and Problem Resolution * The Senior Manager is accountable to guarantee both passive and active monitoring tools are in place and fully functional 24x7x365 to maintain 100 ...
... Incident Management, and Problem Resolution * The Senior Manager is accountable to guarantee both passive and active monitoring tools are in place and fully functional 24x7x365 to maintain 100 ...
Demonstrated proficiency in Incident, Problem, and Change Management (INC/PRB/CHG) - understanding ... how these disciplines interconnect and how to leverage them to drive resiliency outcomes. Knowledge ...
Demonstrated proficiency in Incident, Problem, and Change Management (INC/PRB/CHG) - understanding ... how these disciplines interconnect and how to leverage them to drive resiliency outcomes. Knowledge ...
Collaborates effectively with engineering, operations, security, and vendor teams to support high-impact issues, follows established incident, problem, and change management processes, shares ...
Collaborates effectively with engineering, operations, security, and vendor teams to support high-impact issues, follows established incident, problem, and change management processes, shares ...
Incident Problem Manager information
How does an Incident Problem Manager typically collaborate with technical and non-technical teams during major incidents?
What is the difference between Incident Problem Manager vs Incident Coordinator?
| Aspect | Incident Problem Manager | Incident Coordinator |
|---|---|---|
| Primary Role | Manages the lifecycle of incidents and problems to minimize impact and prevent recurrence | Coordinates incident response activities, ensuring timely resolution and communication |
| Certifications | ITIL Foundation, Problem Management certifications | ITIL Foundation, Incident Management certifications |
| Work Environment | Typically in IT service management teams, focusing on problem analysis | Operational teams, focusing on incident handling and communication |
While both roles are involved in incident management, the Incident Problem Manager focuses on identifying root causes and preventing future issues, whereas the Incident Coordinator handles day-to-day incident response and communication. Both roles are essential for effective IT service delivery but differ in scope and responsibilities.
What does an Incident Problem Manager do?
What are the key skills and qualifications needed to thrive as an Incident Problem Manager, and why are they important?

Full-time
Re-posted 4 days ago
Job description
Purpose of Job
The Senior AI Platform Operations Engineer is accountable for the reliability, operability, and controlled enablement of the organization's AI platform.Â
This role ensures that AI Platform services and solutions are production-ready, secure, observable, and compliant by executing disciplined operational practices, across platform management, monitoring, incident coordination, and governance control enforcement.Â
The incumbent plays a key role in enabling the safe and scalable adoption of AI by ensuring that AI solutions are deployed, monitored, supported, and continuously improved in line with enterprise standards for reliability, security, and compliance
AI Platform Reliability and Operations
  Administer and operate the AI platform to ensure availability, performance, and resilience across environments, integrations, and supporting infrastructure.
  Monitor platform health using dashboards, logs, metrics, and alerts, and coordinate incident and service restoration activities.
  Lead operational triage, escalation coordination, and post-incident reviews to strengthen services stability and resilience.Â
  Track and report on service reliability indicators, incident trends, and operational performance.
AI Platform Enablement & Production Readiness
  Enable approved AI use cases into production by ensuring:Â
o   Environment readiness,Â
o   Dependency validation,Â
o   Completion of operational readiness checklists,Â
o   Structured service transition activities
  Support platform lifecycle management through:
o   Release coordination,Â
o   Change readiness validation,Â
o   Maintenance and capacity planning.
  Ensure AI platform changes meet defined operational and control readiness criteria prior to release
Observability, Automation & AI Ops
  Implement and maintain observability capabilities, including telemetry, logging, metrics, and traces required for enterprise AI operations.
  Analyze operational data to identify anomalies, recurring issues, root-cause patterns.
  Implement AI Ops use cases such as:
o   Alert correlation,Â
o   Anomaly detection,Â
o   Root-cause support,Â
o   Forecasting and predictive insights,Â
o   Automation of repetitive operational tasks.
  Continuously improve operational efficiency through targeted automation and process optimization.
Governance, Risk, & Control Execution
  Execute governance controls for AI solutions, including:
o   Usage and access controls,Â
o   Data privacy considerations,Â
o   Auditability and traceability,Â
o   Human oversight requirements
  Ensure operational practices align with enterprise security policies, risk controls, and compliance requirements.
  Maintain documentation and evidence required for audit, governance reviews, production readiness checkpoints, and control validation.
  Identify control gaps and escalate risks appropriately to relevant governance and risk stakeholders.
AI Asset Visibility & Operational Integrity
  Maintain operational visibility of AI platform assets required for monitoring, support, and cost alignment.Â
  Validate asset ownership, relationships, and lifecycle status in collaboration with application and platform owners.Â
  Support ongoing audits to ensure AI assets and associated cost attribution remain accurate and current.
  5-7 years of experience in platform operations, site reliability engineering, DevOps, cloud operations, or enterprise IT operations.
  Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting.
Technical Expertise:
  Experience with cloud platforms, observability, automation, configuration management, and integration patterns, including Azure Automation runbooks (PowerShell/Python), Azure AI, Copilot integrations, AKS, virtual networks (hub-and-spoke), and App Service.
  Expertise with observability tools such as Azure Monitor, Application Insights, and Grafana.
  Experience with CI/CD and automation tools such as Azure DevOps, GitHub Actions, and Logic Apps.
  Knowledge of configuration management and infrastructure-as-code tools such as Bicep, Terraform, Azure Policy, Key Vault, and relevant open-source technologies.
  Knowledge of integration and event-driven technologies such as API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka.
  Working knowledge of platform-supporting data and search services such as Elastic, Azure AI Search, and Cosmos DB.
  Knowledge of enterprise network, edge security, and related internal platforms such as DNA, Fortinet, and Akamai is an asset.
Additional Capabilities:
  Working knowledge of AI/ML operational concepts, including model lifecycle support, telemetry, governance controls, human-in-the-loop practices, and production monitoring.
  Strong understanding of ITIL/ITSM processes, including change, release, incident, problem, configuration, and service reporting practices.
  Analytical and structured thinker with strong troubleshooting, root-cause analysis, prioritization, and continuous improvement skills.
  Strong service orientation, professional maturity, and the ability to collaborate effectively across operations, engineering, security, risk, data, and business teams.
  Experience creating technical documentation, operational procedures, support playbooks, dashboards, and user guidance materials.
  Knowledge of security, privacy, audit, and compliance considerations relevant to enterprise AI and platform operations.
Job Complexities / Thinking Challenges
This role requires balancing platform reliability, operational efficiency, and governance discipline in a rapidly evolving AI environment.Â