1

Reliability Engineer Manager Jobs in Pennsylvania

Lead, Site Reliability Engineer

Pittsburgh, PA · On-site

$55.25 - $73.50/hr

Define, implement, and operationalize reliability metrics by establishing and managing SLIs, SLOs, and error budgets to quantify and continuously improve service reliability, supporting engineering ...

Lead, Site Reliability Engineer

Pittsburgh, PA · On-site

$55.25 - $73.50/hr

Define, implement, and operationalize reliability metrics by establishing and managing SLIs, SLOs, and error budgets to quantify and continuously improve service reliability, supporting engineering ...

Lead, Site Reliability Engineer

Pittsburgh, PA · On-site

$55.25 - $73.50/hr

Define, implement, and operationalize reliability metrics by establishing and managing SLIs, SLOs, and error budgets to quantify and continuously improve service reliability, supporting engineering ...

Site Reliability Engineer

Malvern, PA · On-site

$56 - $74.25/hr

#W2 Role Senior Reliability Engineer As a Senior Reliability Engineer, you will play a critical role ... Core Responsibilities Team is focused on automating incident response and infrastructure management.

Site Reliability Engineer - Networking

Allentown, PA · On-site

$56.25 - $74.75/hr

As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and ... Design, implementation and management of an overlay network to support 1000's of containers.

A key responsibility is to manage the plant's asset portfolio through a structured Long-Term Asset ... Provide reliability engineering support to Capex projects from early concept through commissioning.

A key responsibility is to manage the plant's asset portfolio through a structured Long-Term Asset ... Provide reliability engineering support to Capex projects from early concept through commissioning.

Software Engineer (.NET & SRE)

Pittsburgh, PA

$55.25 - $73.50/hr

Implement and manage AWS services, including EC2, Lambda, S3, RDS, and others, to support ... Apply SRE principles to enhance system reliability, including monitoring, incident response, and ...

Engineering Manager SRE - Reading, PA

Reading, PA · On-site

$54.75 - $72.75/hr

Key Responsibilites As SRE (Site Responsible Engineer) * Champion the Process Safety Management Program and ensure compliance with OSHA and AkzoNobel standards. * Evaluate and execute continuous ...

Showing results 41-60

Reliability Engineer Manager information

What does a reliability engineer manager do?

A Reliability Engineer Manager oversees teams responsible for improving the reliability and performance of systems, machinery, or processes within an organization. They develop maintenance strategies, lead root cause analyses of failures, and implement best practices to minimize downtime and costs. Additionally, they collaborate with other departments to ensure that reliability goals align with business objectives and compliance standards. Their role is crucial in industries such as manufacturing, energy, and technology, where system uptime and safety are critical.

What are some common challenges reliability engineer managers face when balancing long-term reliability improvements with immediate operational demands?

Reliability Engineer Managers often need to prioritize urgent maintenance issues while also driving long-term reliability initiatives. Balancing these competing demands can be challenging, as immediate equipment failures may require quick fixes that temporarily interrupt ongoing improvement projects. Effective managers work closely with operations, maintenance, and engineering teams to communicate priorities, allocate resources, and implement sustainable solutions that address root causes rather than just symptoms. This role typically involves using data-driven decision-making and fostering a culture of proactive maintenance and continuous improvement.

What are the key skills and qualifications needed to thrive as a reliability engineer manager?

To thrive as a Reliability Engineer Manager, you need a strong background in engineering principles, reliability analysis, and maintenance strategies, typically supported by a degree in engineering and experience in reliability roles. Familiarity with reliability-centered maintenance (RCM), failure mode and effects analysis (FMEA), and asset management software such as SAP or Maximo is common, along with certifications like Certified Reliability Engineer (CRE). Leadership, problem-solving, and effective communication are vital soft skills for managing teams and driving cross-functional initiatives. These competencies are crucial for minimizing downtime, optimizing equipment performance, and ensuring long-term operational efficiency.

What is the difference between Reliability Engineer Manager vs Reliability Engineer?

AspectReliability EngineerReliability Engineer Manager
Required CredentialsBachelor's in Engineering or related field; certifications like CRC, CRESame as Reliability Engineer, plus leadership experience
Work EnvironmentDesign, analyze, and improve system reliability; often in teamsOversees Reliability Engineers; manages projects and teams
Employer & Industry UsageManufacturing, aerospace, energy, automotiveSame industries, with added managerial responsibilities
Common Search & ComparisonFocuses on technical skills and hands-on reliability tasksFocuses on leadership, team management, and strategic planning

The main difference between a Reliability Engineer and a Reliability Engineer Manager lies in their responsibilities. The Reliability Engineer focuses on technical analysis and system improvements, while the Reliability Engineer Manager oversees teams, manages projects, and develops strategies to enhance reliability across the organization.

What are the most commonly searched types of Reliability Engineer jobs in Pennsylvania? The most popular types of Reliability Engineer jobs in Pennsylvania are:
What job categories do people searching Reliability Engineer Manager jobs in Pennsylvania look for? The top searched job categories for Reliability Engineer Manager jobs in Pennsylvania are:
What cities in Pennsylvania are hiring for Reliability Engineer Manager jobs? Cities in Pennsylvania with the most Reliability Engineer Manager job openings:
Infographic showing various Reliability Engineer Manager job openings in Pennsylvania as of August 2026, with employment types broken down into 75% Full Time, and 25% Contract. Highlights an 100% In-person job distribution.

Lead, Site Reliability Engineer

CardWorks

Pittsburgh, PA • On-site

$55.25 - $73.50/hr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

Re-posted yesterday


CardWorks rating

9.1

Company rating: 9.1 out of 10

Based on 8 frontline employees who took The Breakroom Quiz

1st of 21 rated payment service providers


Job description

Become an everyday champion - and build a career where your impact fuels financial progress.


What We Do

CardWorks Financial Group is a diversified financial services platform building ethical solutions across credit, lending, and the full customer lifecycle. Through our family of companies, CardWorks Financial Group tackles the complex challenges that larger financial institutions leave behind. We're embedded throughout the credit card ecosystem as a lender, servicer, and merchant acquirer.

Who We Are

  • Merrick Bank: The bank that builds
  • CardWorks Servicing: One partner, total performance
  • Carson Smithfield: Resolution with respect

With nearly 40 years of operating history, our track record is solid: disciplined in downturns and built to accelerate in recovery. The CardWorks Financial Group companies take precise approach in complex markets, as a top three non-prime focused general purpose card issuer and a top fifteen U.S. merchant acquirer.

Our team tackles the industry's most complex credit and payment challenges. And we believe that excellent work starts with a team that feels supported, respected, and empowered to grow.

CardWorks Servicing, LLC provides end-to end operational servicing functions for credit cards, secured cards, and installment loans. We service consumer and small business loans across the credit spectrum and offers backup servicing and due diligence services to capital providers and trustees.


Founded in 1997, Merrick Bank is an FDIC-insured financial institution headquartered in South Jordan, Utah, with over $10 billion in assets. A wholly owned subsidiary of CardWorks Financial Group, Merrick Bank serves roughly five million cardmembers and more than 100,000 merchant customers, offering credit cards, recreational loans, deposit accounts, merchant services and bank sponsorships to consumers and businesses.

Carson Smithfield, LLC provides a variety of post-charge-off debt recovery services, including digital self-service, IVR, live agent, and external agency management.

Essential Functions:

  • Establish the SRE operating model (service onboarding, engagement model, governance, reliability reviews, production readiness standards, and quarterly planning) and ensure it is adopted across teams.

  • Identify, pilot, and operationalize AI-enabled reliability use cases (e.g., alert noise reduction, incident summarization, correlation/root-cause hypothesis generation, runbook assistance, and auto-remediation with human approval) with appropriate guardrails.

  • Define, implement, and operationalize reliability metrics by establishing and managing SLIs, SLOs, and error budgets to quantify and continuously improve service reliability, supporting engineering and business decisions.

  • Own the centralized SRE service engagement model by defining service tiers, onboarding criteria, reliability standards, and a transparent intake/prioritization process aligned to business criticality.

  • Define and enforce error budget policies (including escalation paths and release risk decisions) in partnership with Product/Engineering, using SLO attainment to guide trade-offs between feature velocity and reliability

  • Establish and maintain centralized "paved road" reliability standards and assets (instrumentation conventions, golden signals, alerting standards, runbook templates, SLO dashboards) that product teams can adopt with minimal friction.

  • Design the on-call and escalation model for a centralized SRE team (e.g., SRE overlay for major incidents, defined handoffs with service owners, and clear ownership boundaries) to improve response quality without creating single-team dependency.

  • Design and engineer automation and observability solutions by developing tooling, dashboards, and systems to reduce operational toil (measure, report, and drive toil down over time), enhance system visibility, and accelerate delivery.

  • Participates in incident and problem management by serving as incident coordinator for high-severity events, driving cross-functional responses, conducting blameless root cause analysis, running post-incident reviews (postmortems) with clear owners and due dates, ensuring remedial actions drive reliability improvements.

  • Oversee operational readiness and performance by managing capacity planning, validating disaster recovery, conducting production readiness reviews, and ensuring systems meet availability, scalability, and recovery expectations.

  • Partner with security, risk, and compliance teams to align reliability goals with governance and compliance requirements, ensuring secure, auditable, and well-documented practices.

  • Collaborate across the organization by working closely with end users, product management, development, architecture, and IT Operational teams to embed reliability principles throughout the software development lifecycle, including service onboarding, reliability reviews, and shared SLO ownership.

  • Champion reliability as a core product feature by promoting reliability throughout all phases of development, advocating for continuous improvement, and communicating key metrics and potential customer impact to stakeholders.

  • Train, mentor, and upskill engineering teams by coaching engineers in SRE practices, supporting junior team members, and fostering a culture of shared ownership and accountability for reliability, including influencing teams without direct authority through standards, data, and executive-aligned priorities. Remain current on the latest SRE trends and best practices, including observability, AI-enabled operations (AIOps), and SLO management, and implement these methodologies to effectively support desired business outcomes. Evaluate AI tools for reliability with security/privacy/compliance guardrails (e.g., data handling, prompt/content controls, auditability) and measure impact.

  • Participate in on-call rotations and operational support for SRE-supported systems and products.

Summary of Qualifications:

  • Experience in Site Reliability Engineering with a track record of delivering measurable improvements in uptime, scalability, release stability, and overall reliability in complex enterprise environments.

  • Demonstrated experience standing up or significantly maturing an SRE practice (operating model, SRE/service engagement, production readiness, incident/postmortem program, and reliability roadmap).

  • Hands-on experience applying AI/ML to operations (AIOps) or GenAI in production support workflows, with a focus on measurable outcomes (MTTD/MTTR, alert fatigue reduction, change failure rate) and responsible use controls.

  • Proven ability to establish Service Level Indicators (SLIs) and SLOs in production environments, including hands-on definition and implementation.

  • Demonstrated background in production incident response, leading resolution efforts, conducting blameless post-incident reviews, and implementing actionable remediation strategies.

  • Strong observability and telemetry expertise in designing instrumentation, building actionable dashboards and alerts, and delivering proactive reliability insights using metrics, logs, and traces.

  • Infrastructure engineering experience with strong Infrastructure as Code skills using tools such as Terraform and Ansible.

  • Thorough understanding and practical experience in CI/CD pipeline design, optimization, and troubleshooting using modern tooling and platforms such as Azure DevOps, GitHub Actions, Jenkins, or GitLab CI, with an emphasis on speed, reliability, and security.

  • Practical knowledge of containerization and platform modernization, including architecting and operating containerized workloads with Docker, VMware, and Kubernetes (or comparable orchestration platforms) to modernize legacy applications and improve fault tolerance.

  • Knowledge of emerging reliability practices, including SLO automation platforms, AIOps, or predictive operations to advance proactive reliability management.

  • Preferred certifications include AWS Professional, Terraform, Ansible, Azure DevOps, Octopus Deploy or other automation-focused credentials that demonstrate continuous technical development.

Education and Experience:

  • Master's degree in computer science, Engineering, or equivalent practical experience designing and operating production systems at scale.

  • 7+ years of experience in Site Reliability Engineering.

Ideally, the qualified candidate will work at the following location(s): Woodbury, NY; Pittsburgh, PA, Orlando, Fl, South Jordan, UT. A hybrid work model or fully remote model can be considered based on hiring manager decision and priorities of the role.

The salary range for this position, if located in NY Metro/NY State is $146,032 to $162,257. However, please note that the salary range will vary for other geographic areas.

#INDHP

Our Employee Value Proposition

  • Competitive Pay, including a Bonus Target or Variable Pay Incentive Program
  • Benefits Package -Medical, Dental, and Vision (plus much more)
  • 401(k) Plan with Company Match
  • Short- & Long-Term Disability
  • Wellness Programs
  • Group Life and AD&D Insurance
  • Paid Vacation, Sick Days and bank Holidays
  • Employee Engagement Activities including Employee Appreciation Day, DEI Employee Resource Groups, Corporate Social Responsibility, Service Recognition

We offer a total rewards package comprised of a competitive base rate of pay, variable pay incentive programs based on the role, and a comprehensive benefit suite. Offered rates of pay are determined based on job-related knowledge, relevant experience, skills, certifications, and geographic location.


We are proud to be an equal opportunity employer. All qualified applicants will receive consideration without regard to age, race, color, sex, or gender identity/expression (including pregnancy, childbirth, transgender status, or sexual orientation), religion or creed, ancestry, citizenship, national origin, disability, military or veteran status, marital status, genetic information, or any other characteristic protected by applicable law.

We do not tolerate discrimination, harassment, or retaliation. Employment decisions are based solely on qualifications, merit, and business needs. Everyone is welcome here, and we hire based on your ability to do the job, not any protected characteristics.

If you need help or reasonable accommodation during the application or hiring process, please let your TA Partner know.



What CardWorks employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom