1

Incident Problem Manager Jobs in Pennsylvania (NOW HIRING)

Site Reliability Engineer

Pittsburgh, PA ยท On-site

$55.25 - $73.50/hr

Incident & Problem Management * Linux * Windows Server * Oracle / PL/SQL / DB2 * Dynatrace / DT Managed * GlassBox / ITCAM / TrueSight / OEM * Tomcat / Apache / WebSphere (WAS) / IIS * REST & SOAP ...

Site Reliability Engineer

Pittsburgh, PA ยท On-site

$53.25 - $70.75/hr

Incident & Problem Management * Linux * Windows Server * Oracle / PL/SQL / DB2 * Dynatrace / DT Managed * GlassBox / ITCAM / TrueSight / OEM * Tomcat / Apache / WebSphere (WAS) / IIS * REST & SOAP ...

Incident & Problem Management: Lead all Major Incidents and drive root-cause analysis to minimize major and recurring incidents, turning incident data into lasting fixes. * Change Management: Run the ...

ServiceNow Architect

Philadelphia, PA ยท On-site

$54.50 - $75/hr

Experience in implementing ITIL processes (Incident, Problem, Change Management, etc.) in ServiceNow ITOM : Experience in multiple End-to-End implementation of ITOM including CMDB processes (Event ...

ServiceNow Architect

Philadelphia, PA ยท On-site

$54.50 - $75/hr

Experience in implementing ITIL processes (Incident, Problem, Change Management, etc.) in ServiceNow ITOM : Experience in multiple End-to-End implementation of ITOM including CMDB processes (Event ...

Showing results 21-40

Incident Problem Manager information

What does an Incident Problem Manager do?

An Incident Problem Manager is responsible for overseeing the process of identifying, investigating, and resolving incidents and underlying problems within an organization's IT systems. They work to minimize the impact of disruptions, coordinate responses to incidents, and analyze root causes to prevent future occurrences. This role often involves collaborating with technical teams, managing communication with stakeholders, and ensuring that procedures are followed according to IT service management frameworks like ITIL. Their goal is to improve IT service reliability and reduce downtime for the business.

How does an Incident Problem Manager typically collaborate with technical and non-technical teams during major incidents?

Incident Problem Managers play a critical role in bridging communication between technical teams (like IT support, network engineers, or developers) and non-technical stakeholders (such as business unit leaders or customer service). During major incidents, they coordinate response efforts, facilitate status updates, and ensure all parties are aligned on next steps and remediation plans. Effective collaboration involves translating complex technical issues into clear, actionable information for non-technical audiences, managing expectations, and driving post-incident reviews to prevent recurrence. This cross-functional coordination is essential for minimizing business impact and ensuring swift resolution.

What are the key skills and qualifications needed to thrive as an Incident Problem Manager, and why are they important?

To thrive as an Incident Problem Manager, you need strong analytical skills, IT service management knowledge, and experience with incident and problem resolution processes, often supported by ITIL certification. Familiarity with ITSM tools like ServiceNow, Jira Service Management, or BMC Remedy is typically required. Exceptional communication, leadership, and critical thinking abilities enable effective coordination and root-cause analysis across teams. These skills are crucial to minimizing downtime, ensuring service continuity, and driving long-term improvements in IT operations.

What is the difference between Incident Problem Manager vs Incident Coordinator?

AspectIncident Problem ManagerIncident Coordinator
Primary RoleManages the lifecycle of incidents and problems to minimize impact and prevent recurrenceCoordinates incident response activities, ensuring timely resolution and communication
CertificationsITIL Foundation, Problem Management certificationsITIL Foundation, Incident Management certifications
Work EnvironmentTypically in IT service management teams, focusing on problem analysisOperational teams, focusing on incident handling and communication

While both roles are involved in incident management, the Incident Problem Manager focuses on identifying root causes and preventing future issues, whereas the Incident Coordinator handles day-to-day incident response and communication. Both roles are essential for effective IT service delivery but differ in scope and responsibilities.

What are popular job titles related to Incident Problem Manager jobs in Pennsylvania?

For Incident Problem Manager jobs in Pennsylvania, the most frequently searched job titles are:

What job categories do people searching Incident Problem Manager jobs in Pennsylvania look for?

The top searched job categories for Incident Problem Manager jobs in Pennsylvania are:

What cities in Pennsylvania are hiring for Incident Problem Manager jobs?

Cities in Pennsylvania with the most Incident Problem Manager job openings:

Infographic showing various Incident Problem Manager job openings in Pennsylvania as of August 2026, with employment types broken down into 82% Full Time, 17% Part Time, and 1% Contract. Highlights an 81% Physical, 2% Hybrid, and 17% Remote job distribution.

Site Reliability Engineer

System One

Pittsburgh, PA โ€ข On-site

$55.25 - $73.50/hr

Full-time

Posted 22 days ago


Job description

Senior Site Reliability Engineer (SRE) Location: Pittsburgh, PA / Cleveland, OH / Dallas, TX FTE Position Overview We are seeking an experienced Senior Site Reliability Engineer (SRE) to support production operations, application reliability, performance management, and continuous improvement initiatives. The selected candidate will work closely with production support and engineering teams to ensure critical internal and external applications maintain appropriate levels of availability, reliability, and uptime. This role requires strong experience in production support, incident management, monitoring, troubleshooting, log analysis, automation identification, infrastructure technologies, databases, and application servers. The SRE will also provide technical leadership and collaborate with geographically distributed teams. Key Skills

  • Site Reliability Engineering (SRE)
  • Production Support / Application Support
  • Incident & Problem Management
  • Linux
  • Windows Server
  • Oracle / PL/SQL / DB2
  • Dynatrace / DT Managed
  • GlassBox / ITCAM / TrueSight / OEM
  • Tomcat / Apache / WebSphere (WAS) / IIS
  • REST & SOAP Web Services
  • Log Analysis & Troubleshooting
  • AIOps / NLP
  • Monitoring & Performance Management
  • Automation
  • Root Cause Analysis
  • Business Analytics
  • Agile
  • Technical Leadership
  • Client-Facing Production Support
Responsibilities
  • Monitor distributed systems and proactively identify potential production issues.
  • Support troubleshooting and participate in on-call activities.
  • Manage, track, and coordinate production incidents and application outages.
  • Lead incident-analysis and problem-management meetings.
  • Identify opportunities for operational and production-support automation.
  • Monitor applications and related infrastructure to maintain system reliability.
  • Coordinate follow-up activities through incident resolution and closure.
  • Troubleshoot complex application issues using system and application logs.
  • Participate in critical incident calls and contribute technical expertise toward resolution.
  • Perform root cause analysis and recommend corrective actions.
  • Research and reproduce user issues to validate solutions.
  • Resolve technical problems that cannot be handled by junior team members.
  • Provide technical guidance and solutions to the production-support team.
  • Introduce process improvements and innovative solutions for operational challenges.
  • Develop and maintain SOPs, operational procedures, and knowledge documentation.
  • Collaborate with offshore and geographically distributed teams.
  • Work with client technical teams, SMEs, and leadership.
  • Support extended or weekend hours when required during critical production events.
  • Participate in overlapping business-hour shifts for critical meetings and activities.
Required Qualifications
  • 5+ years of overall IT experience.
  • 2–3 years of business analytics and technical leadership experience.
  • Strong experience with production/application support in a client-facing environment.
  • Strong understanding of Site Reliability Engineering and production operations.
  • Hands-on experience troubleshooting production applications and analyzing log files.
  • Strong knowledge of system-management, monitoring, and support analytics tools.
  • Experience with incident management, root cause analysis, and problem resolution.
  • Strong understanding of AIOps and NLP concepts.
  • Experience identifying opportunities for automation and process improvement.
  • Strong problem-solving and analytical capabilities.
  • Ability to recommend efficient and cost-effective technical solutions.
  • Experience working with geographically distributed/onshore-offshore teams.
  • Excellent client-facing verbal and written communication skills.
Database Technologies Strong knowledge of:
  • Oracle
  • PL/SQL
  • DB2
Web Services Experience developing and consuming:
  • REST APIs
  • SOAP Web Services
Experience should preferably be within an operational/production environment. Application Servers / Web Servers Strong knowledge of:
  • Tomcat
  • Apache
  • WebSphere (WAS)
  • IIS
Operating Systems
  • Extensive experience with Linux
  • Good understanding of Windows Server
  • Linux and Windows server configuration and troubleshooting
Monitoring & Support Tools Experience with monitoring tools such as:
  • Dynatrace
  • Dynatrace Managed / DT Managed
  • GlassBox
  • ITCAM / ITCAMS
  • TrueSight
  • Oracle Enterprise Manager (OEM)
Additional Skills
  • Agile methodology
  • SOP and technical documentation
  • Performance management
  • System reliability and availability
  • Production incident coordination
  • Technical research and solution evaluation
  • Process improvement
  • Automation opportunity identification
  • Strong stakeholder and client communication

#M1 #DI-CB2 #L1 - KB1

Ref: #404-IT Pittsburgh


System One logo

About System One

Sourced by ZipRecruiter

System One helps employers get work done more efficiently and economically without compromising quality. Over our 35+ year history, we've helped connect thousands of talented people with innovative companies. The excitement of a perfect fit motivates us every single day.

Industry

Business consulting services and recruiting and staffing services

Company size

5,001 - 10,000 Employees

Headquarters location

Pittsburgh, PA, US