1

Observability Site Reliability Engineer Jobs in Pennsylvania

Lead, Site Reliability Engineer

Pittsburgh, PA · On-site

$55.25 - $73.50/hr

Remain current on the latest SRE trends and best practices, including observability, AI-enabled operations (AIOps), and SLO management, and implement these methodologies to effectively support ...

$48.75 - $64.75/hr

As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable ... Implement and improve monitoring, alerting, logging, tracing, and observability capabilities across ...

$48.75 - $64.75/hr

As a Senior Site Reliability Engineer at Autodesk, you can help us build and operate reliable ... Implement and improve monitoring, alerting, logging, tracing, and observability capabilities across ...

Staff Site Reliability Engineer

Crum Lynne, PA

$54.50 - $72.25/hr

The Site Reliability Engineering team drives reliability strategy, elevates engineering standards ... Expert-level command of monitoring, observability, and alerting platforms (e.g., Datadog ...

Staff Site Reliability Engineer

Crum Lynne, PA · On-site

$54.50 - $72.25/hr

The Site Reliability Engineering team drives reliability strategy, elevates engineering standards ... Expert-level command of monitoring, observability, and alerting platforms (e.g., Datadog ...

Site Reliability Engineer

Pittsburgh, PA · On-site

$55.25 - $73.50/hr

Provide SRE production support for Mainframe applications, APIs, Web Services, and Microservices ... Build and enhance monitoring, alerting, dashboards, observability, and self-healing capabilities.

Site Reliability Engineer

Pittsburgh, PA · On-site

$55.25 - $73.50/hr

Senior Site Reliability Engineer (SRE) Location: Pittsburgh, PA / Cleveland, OH / Dallas, TX FTE Position Overview We are seeking an experienced Senior Site Reliability Engineer (SRE) to support ...

Site Reliability Engineer

Pittsburgh, PA · On-site

$53.25 - $70.75/hr

Senior Site Reliability Engineer (SRE) Location: Pittsburgh, PA / Cleveland, OH / Dallas, TX FTE Position Overview We are seeking an experienced Senior Site Reliability Engineer (SRE) to support ...

next page

Showing results 1-20

Observability Site Reliability Engineer information

What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?

AspectObservability Site Reliability EngineerMonitoring Engineer
FocusEnsuring system reliability through observability, automation, and incident responseImplementing and managing monitoring tools and dashboards
SkillsCloud platforms, scripting, incident management, observability toolsMonitoring tools, alerting systems, data analysis
Work EnvironmentDevOps teams, cloud infrastructure, large-scale systemsOperations teams, infrastructure monitoring

While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.

What are popular job titles related to Observability Site Reliability Engineer jobs in Pennsylvania?

For Observability Site Reliability Engineer jobs in Pennsylvania, the most frequently searched job titles are:

What job categories do people searching Observability Site Reliability Engineer jobs in Pennsylvania look for?

The top searched job categories for Observability Site Reliability Engineer jobs in Pennsylvania are:

What cities in Pennsylvania are hiring for Observability Site Reliability Engineer jobs?

Cities in Pennsylvania with the most Observability Site Reliability Engineer job openings:

Lead, Site Reliability Engineer

Pittsburgh, PA • On-site


CardWorks

9.1

Company rating: 9.1 out of 10

Based on 8 frontline employees who took The Breakroom Quiz

1st of 21 rated payment service providers

Great coworkers

People enjoy working here

Good employer


$55.25 - $73.50/hr

Full-time

This job post has expired 3 days ago. Applications are no longer accepted.


Job description

Job Summary:
CardWorks is a diversified financial services platform that provides ethical solutions across credit and lending. They are seeking a Lead Site Reliability Engineer to establish and operationalize the SRE operating model, focusing on improving service reliability and integrating AI-enabled solutions into their processes.
Responsibilities:
• Establish the SRE operating model (service onboarding, engagement model, governance, reliability reviews, production readiness standards, and quarterly planning) and ensure it is adopted across teams.
• Identify, pilot, and operationalize AI-enabled reliability use cases (e.g., alert noise reduction, incident summarization, correlation/root-cause hypothesis generation, runbook assistance, and auto-remediation with human approval) with appropriate guardrails.
• Define, implement, and operationalize reliability metrics by establishing and managing SLIs, SLOs, and error budgets to quantify and continuously improve service reliability, supporting engineering and business decisions.
• Own the centralized SRE service engagement model by defining service tiers, onboarding criteria, reliability standards, and a transparent intake/prioritization process aligned to business criticality.
• Define and enforce error budget policies (including escalation paths and release risk decisions) in partnership with Product/Engineering, using SLO attainment to guide trade-offs between feature velocity and reliability.
• Establish and maintain centralized “paved road” reliability standards and assets (instrumentation conventions, golden signals, alerting standards, runbook templates, SLO dashboards) that product teams can adopt with minimal friction.
• Design the on-call and escalation model for a centralized SRE team (e.g., SRE overlay for major incidents, defined handoffs with service owners, and clear ownership boundaries) to improve response quality without creating single-team dependency.
• Design and engineer automation and observability solutions by developing tooling, dashboards, and systems to reduce operational toil (measure, report, and drive toil down over time), enhance system visibility, and accelerate delivery.
• Participates in incident and problem management by serving as incident coordinator for high-severity events, driving cross-functional responses, conducting blameless root cause analysis, running post-incident reviews (postmortems) with clear owners and due dates, ensuring remedial actions drive reliability improvements.
• Oversee operational readiness and performance by managing capacity planning, validating disaster recovery, conducting production readiness reviews, and ensuring systems meet availability, scalability, and recovery expectations.
• Partner with security, risk, and compliance teams to align reliability goals with governance and compliance requirements, ensuring secure, auditable, and well-documented practices.
• Collaborate across the organization by working closely with end users, product management, development, architecture, and IT Operational teams to embed reliability principles throughout the software development lifecycle, including service onboarding, reliability reviews, and shared SLO ownership.
• Champion reliability as a core product feature by promoting reliability throughout all phases of development, advocating for continuous improvement, and communicating key metrics and potential customer impact to stakeholders.
• Train, mentor, and upskill engineering teams by coaching engineers in SRE practices, supporting junior team members, and fostering a culture of shared ownership and accountability for reliability, including influencing teams without direct authority through standards, data, and executive-aligned priorities. Remain current on the latest SRE trends and best practices, including observability, AI-enabled operations (AIOps), and SLO management, and implement these methodologies to effectively support desired business outcomes. Evaluate AI tools for reliability with security/privacy/compliance guardrails (e.g., data handling, prompt/content controls, auditability) and measure impact.
• Participate in on-call rotations and operational support for SRE-supported systems and products.
Qualifications:
Required:
• Experience in Site Reliability Engineering with a track record of delivering measurable improvements in uptime, scalability, release stability, and overall reliability in complex enterprise environments.
• Demonstrated experience standing up or significantly maturing an SRE practice (operating model, SRE/service engagement, production readiness, incident/postmortem program, and reliability roadmap).
• Hands-on experience applying AI/ML to operations (AIOps) or GenAI in production support workflows, with a focus on measurable outcomes (MTTD/MTTR, alert fatigue reduction, change failure rate) and responsible use controls.
• Proven ability to establish Service Level Indicators (SLIs) and SLOs in production environments, including hands-on definition and implementation.
• Demonstrated background in production incident response, leading resolution efforts, conducting blameless post-incident reviews, and implementing actionable remediation strategies.
• Strong observability and telemetry expertise in designing instrumentation, building actionable dashboards and alerts, and delivering proactive reliability insights using metrics, logs, and traces.
• Infrastructure engineering experience with strong Infrastructure as Code skills using tools such as Terraform and Ansible.
• Thorough understanding and practical experience in CI/CD pipeline design, optimization, and troubleshooting using modern tooling and platforms such as Azure DevOps, GitHub Actions, Jenkins, or GitLab CI, with an emphasis on speed, reliability, and security.
• Practical knowledge of containerization and platform modernization, including architecting and operating containerized workloads with Docker, VMware, and Kubernetes (or comparable orchestration platforms) to modernize legacy applications and improve fault tolerance.
• Knowledge of emerging reliability practices, including SLO automation platforms, AIOps, or predictive operations to advance proactive reliability management.
• Master’s degree in computer science, Engineering, or equivalent practical experience designing and operating production systems at scale.
• 7+ years of experience in Site Reliability Engineering.
Preferred:
• Preferred certifications include AWS Professional, Terraform, Ansible, Azure DevOps, Octopus Deploy or other automation-focused credentials that demonstrate continuous technical development.
Company:
Cardworks is a service provider provides comprehensive service and support to bankcard issuers. Founded in 1987, the company is headquartered in Woodbury, USA, with a team of 1001-5000 employees. The company is currently Late Stage.

What CardWorks employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom