1

Incident Problem Management Engineer Jobs (NOW HIRING)

... engineering, operations, and business teams. This role also contributes to problem management, change coordination, and operational excellence initiatives, with a primary focus on incident leadership ...

Problem Manager

Boston, MA · Remote

$55 - $65/hr

... Incident management,Itil,Servicenow,Problem management,Troubleshooting,Itsm,Root cause analysis Additional Skills & Qualifications The Problem Management Engineer is responsible for leading and ...

Incident Manager

Jersey City, NJ · On-site

$110K - $130K/yr

... engineering, operations, and business teams. This role also contributes to problem management, change coordination, and operational excellence initiatives, with a primary focus on incident leadership ...

We are engineers, high line workers, power plant managers, accountants, electricians, project ... Develop and implement incident and problem management processes and protocols. * Conduct post ...

NOC Incident Manager

San Diego, CA · On-site

$89K - $116K/yr

Coordinate escalation to L2/L3 engineering, vendors, or facilities teams to accelerate fault ... Partner with Problem Management to identify chronic issues and implement long‑term remediation.

Hold engineering and product teams accountable for root cause identification and fix delivery ... Oversee analysis of incident trends to identify systemic issues * Ensure accurate, consistent ...

Showing results 21-40

Incident Problem Management Engineer information

See salary details

$36.5K

$163.4K

$193.5K

How much do incident problem management engineer jobs pay per year?

As of Sep 9, 2026, the average yearly pay for incident problem management engineer in the United States is $163,404.00, according to ZipRecruiter salary data. Most workers in this role earn between $129,000.00 and $193,000.00 per year, depending on experience, location, and employer.

What is an incident problem management engineer?

An Incident Problem Management Engineer is an IT professional responsible for identifying, analyzing, and resolving incidents and problems within an organization's technology environment. They manage the process of restoring normal service operation as quickly as possible after incidents, and work to identify root causes to prevent future issues. Their role often involves close collaboration with other IT teams, conducting post-incident reviews, and implementing solutions to improve system reliability and performance. They play a crucial role in maintaining business continuity and minimizing the impact of IT disruptions.

How does an incident problem management engineer typically interact with other IT teams during major incidents?

An Incident Problem Management Engineer works closely with cross-functional IT teams, such as network operations, application support, and cybersecurity, to coordinate rapid response during major incidents. They facilitate communication between stakeholders, ensure accurate documentation of issues, and lead root cause analysis sessions post-incident. Strong collaboration and clear communication skills are essential, as the engineer often acts as the central point of contact to drive incident resolution and implement long-term preventative measures.

What key skills and qualifications are needed to thrive as an incident problem management engineer, and why are they important?

To excel as an Incident Problem Management Engineer, you need strong analytical skills, a solid understanding of ITIL processes, and experience with incident and problem resolution in complex IT environments. Familiarity with tools like ServiceNow, Jira, and monitoring systems, as well as ITIL or relevant certifications, is highly valued. Excellent communication, critical thinking, and the ability to remain calm under pressure are standout soft skills for this role. These abilities are crucial for minimizing service disruptions, addressing root causes efficiently, and ensuring continuous improvement in IT service management.

What is the difference between Incident Problem Management Engineer vs Network Operations Center (NOC) Engineer?

AspectIncident Problem Management EngineerNetwork Operations Center (NOC) Engineer
CertificationsITIL, Network+, CCNANetwork+, CCNA, CompTIA Security+
Work EnvironmentIT service management, incident analysis, problem resolutionMonitoring network infrastructure, troubleshooting connectivity issues
Employer & IndustryIT service providers, large enterprisesTelecom, internet service providers, data centers
Common Search & ComparisonFocus on incident and problem resolution processesFocus on network monitoring and troubleshooting

The Incident Problem Management Engineer primarily handles incident analysis and problem resolution within IT service management frameworks, often working on root cause analysis and process improvement. In contrast, the Network Operations Center (NOC) Engineer monitors network infrastructure, troubleshoots connectivity issues, and ensures network uptime. While both roles require technical certifications and involve troubleshooting, their focus areas and work environments differ significantly.

What are popular job titles related to Incident Problem Management Engineer jobs?

For Incident Problem Management Engineer jobs, the most frequently searched job titles are:

Incident and Problem Management Analyst

Atlanta, GA • On-site

Inspire Brands
Food Services and Drinking Places • 10K+ employees

Full-time

Re-posted 21 days ago


Inspire Brands rating

5.8

Company rating: 5.8 out of 10

Based on 56 frontline employees who took The Breakroom Quiz

29th of 108 rated fast food restaurants


Job description

The Incident & Problem Management Analyst is responsible for ensuring a cohesive Incident Management, Major Incident Management, and Problem Management practice across Inspire Brands Shared Services. This role leads the Major Incident Management process supporting Shared Services, E-Commerce, Infrastructure, and other critical enterprise platforms.

The analyst serves as the central coordinator during major incidents, leading technical bridge calls, driving rapid engagement of support teams, managing executive communications, and ensuring timely service restoration. This role is responsible for crafting clear, concise, and business-focused communications for senior leadership throughout the incident lifecycle, translating complex technical issues into actionable business updates.

In addition to Major Incident Management responsibilities, the analyst owns key aspects of the Problem Management practice, driving root cause analysis, coordinating corrective actions, and identifying opportunities for long-term incident reduction through ITIL-aligned best practices.

Success in this role requires the ability to quickly develop and maintain a strong functional understanding of Inspire Brands' critical technology ecosystem, including MDBP, IDP, digital commerce platforms, integrations, APIs, infrastructure components, and end-to-end transaction flows. The analyst must be capable of rapidly assessing business impact, understanding system dependencies, and coordinating technical teams during high-severity incidents.

Responsibilities

Major Incident Management

  • Lead and coordinate Major Incident Management activities for enterprise technology services and platforms.
  • Serve as Incident Commander during high-priority incidents and service disruptions.
  • Facilitate and lead technical bridge calls during critical incidents.
  • Engage and coordinate appropriate technical resources, vendors, and business stakeholders to expedite resolution.
  • Drive accountability across support teams and ensure clear ownership of resolution activities.
  • Maintain command and control throughout the incident lifecycle.
  • Quickly assess business impact and communicate risk, customer impact, and service degradation to stakeholders.
  • Develop and distribute executive-facing communications, including:
    • Initial incident notifications
    • Status updates
    • Business impact assessments
    • Executive summaries
    • Resolution communications
    • Post-incident reports
  • Translate complex technical information into concise, business-focused messaging for senior leadership.
  • Coordinate post-incident reviews and ensure action items are identified and tracked to completion.
  • Support high-visibility business events, hypercare periods, and executive escalations when service disruptions occur.

Problem Management

  • Identify, classify, and manage problem records to prevent recurring incidents.
  • Analyze incident trends and recurring issues to identify opportunities for systemic improvements.
  • Drive Root Cause Analysis (RCA) efforts following major incidents and significant service disruptions.
  • Coordinate internal and external stakeholders to investigate root causes.
  • Ensure corrective actions are assigned, tracked, and completed.
  • Manage problem records through closure while ensuring proper documentation and knowledge transfer.
  • Promote proactive problem management practices across technology teams.
  • Drive continuous service improvement initiatives that reduce operational risk and improve service stability.

Operational Excellence & Service Governance

  • Coordinate resources and processes required to resolve enterprise-wide outages and service disruptions.
  • Support the continual improvement of Incident, Major Incident, and Problem Management processes.
  • Maintain and enhance operational documentation, standard operating procedures, and knowledge articles.
  • Ensure quality procedures and solutions are documented and readily available to support teams.
  • Collaborate with technical and business teams to improve operational maturity and service reliability.
  • Support reporting, service governance, and operational metrics initiatives.
  • Contribute to ITSM maturity improvements aligned to ITIL best practices and organizational objectives.

Technical Understanding & Impact Assessment

  • Develop and maintain a high-level understanding of Inspire Brands' technology ecosystem, including MDBP, IDP, digital platforms, integrations, APIs, and supporting infrastructure.
  • Understand application dependencies, transaction flow paths, integration points, and service relationships.
  • Quickly identify and assess outage impacts across interconnected business systems.
  • Evaluate risks and potential customer impact during service disruptions.
  • Partner with application, infrastructure, network, security, and vendor teams to accelerate incident diagnosis and resolution.
  • Support continual refinement of service dependency mapping and operational readiness processes.

Communication & Stakeholder Management

  • Communicate effectively with business stakeholders, technical teams, vendors, management, and senior leadership.
  • Develop and deliver executive-level incident communications during high-pressure situations.
  • Facilitate collaboration across multiple technical teams and organizational functions.
  • Influence stakeholders and drive accountability toward resolution and corrective action completion.
  • Maintain professionalism and composure during critical incidents and executive escalations.

Qualifications

Education

  • Bachelor's Degree in Computer Science, Information Technology, Management Information Systems, or a related field.

Experience

  • Minimum of 3 years of experience in a medium to large enterprise environment.
  • 3+ years of experience defining, implementing, and improving IT Service Management (ITSM) processes.
  • 3+ years of experience supporting Incident Management, Major Incident Management, and Problem Management practices.
  • Experience leading incident bridge calls and coordinating cross-functional technical teams.
  • Experience facilitating Root Cause Analysis (RCA) and driving corrective actions to resolution.
  • Experience creating executive-level communications during technology incidents and service disruptions

Knowledge, Skills & Abilities

  • Strong understanding of Incident Management and Problem Management principles within an ITIL framework.
  • Knowledge of the Three Lines of Defense model, preferably within a multi-brand shared services organization.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Demonstrated ability to rapidly learn complex enterprise application ecosystems and business processes.
  • Functional understanding of system architecture, application dependencies, APIs, integrations, and end-to-end transaction flows.
  • Ability to assess outage impacts across interconnected business and technology platforms.
  • Exceptional verbal and written communication skills.
  • Proven ability to create concise, business-focused executive communications.
  • Strong facilitation skills with the ability to lead and direct major incident bridge calls.
  • Ability to influence stakeholders at all levels of the organization.
  • Process improvement mindset focused on operational excellence and service reliability.
  • Experience with IT Service Management platforms such as ServiceNow.
  • Ability to perform effectively in high-pressure environments requiring rapid decision-making and coordination. [Re: HSC Ba...tification | Outlook]

Preferred Qualifications

  • ITIL Foundation Certification or equivalent ITSM certification.
  • Experience supporting digital commerce, customer identity, integration, middleware, or enterprise platform environments.
  • Experience supporting highly integrated enterprise ecosystems with complex application dependencies.
  • Familiarity with transaction flow analysis, service dependency mapping, and impact assessment methodologies.
  • Experience supporting multi-brand organizations and shared services operating models.
  • Experience with reporting, service analytics, and continuous improvement programs.

Core Competencies

  • Major Incident Leadership
  • Problem Management
  • Root Cause Analysis
  • Executive Communication
  • Stakeholder Management
  • Operational Excellence
  • Service Governance
  • Process Improvement
  • Continuous Service Improvement
  • Cross-Functional Collaboration
  • Critical Thinking & Decision Making
  • Customer Focus

Why Join Inspire Brands?

The Incident & Problem Management Analyst plays a critical role in maintaining the stability and reliability of technology services that support some of the world's most recognized restaurant brands. This position provides an opportunity to lead critical incident response efforts, influence service reliability improvements, partner with enterprise stakeholders, and contribute directly to operational excellence across Inspire Brands.


Inspire is a multi-brand restaurant company whose portfolio includes more than 33,300 Arby's, Baskin-Robbins, Buffalo Wild Wings, Dunkin', Jimmy John's, and SONIC restaurants worldwide. We're made up of some of the world's most iconic restaurant brands, but we're much more than just a restaurant company. We're a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple-it's an experience. At Inspire, that's our purpose: to ignite and nourish flavorful experiences.

What Inspire Brands employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Inspire Brands logo

About Inspire Brands

Sourced by ZipRecruiter

Inspire Brands Inc., located in Atlanta, GA, United States, operates in the foodservice industry as a multi-brand restaurant company, making it among the biggest restaurant companies globally. Their portfolio includes well-known restaurant brands such as Arby's, Buffalo Wild Wings, Sonic, and Jimmy John's, reflecting their commitment to innovation and quality. Founded in 2018 as a result of a consolidation of various restaurant brands under one corporate umbrella, Inspire Brands was formed with a vision to invigorate excellent brands and supercharge their long-term growth.

Industry

Food services and drinking places

Company size

10,000+ Employees

Headquarters location

Atlanta, GA, US

Year founded

2018