1

Senior Observability Engineer Jobs in Philadelphia, PA

... senior engineers and platform leadership. The Intermediate Platform Engineer partners with ... Contribute to the design, implementation, and support of enterprise monitoring and observability ...

Senior Frontend Engineer

Wilmington, DE ยท On-site +1

$118K - $163K/yr

Senior Front-End Developer Frontend Engineering React โ€ข TypeScript โ€ข TanStack Router โ€ข ... Improve accessibility, testing, observability, and overall developer experience. Contribute to ...

Senior DevOps Engineer

Philadelphia, PA ยท Hybrid

$131K - $168K/yr

Section 1: Position Summary We are seeking a Senior DevOps Engineer with 10+ years of handson ... Define and enforce observability standards across logs, metrics, and traces * Implement service ...

Senior Core Network Engineer

Philadelphia, PA ยท Hybrid

$104K - $143K/yr

This position requires a senior-level engineer who can operate both strategically and tactically ... Drive continuous improvement initiatives focused on resiliency, automation, observability, and ...

Senior Cloud Engineer

Broomall, PA ยท On-site +1

$103K - $142K/yr

We are seeking a hands-on Senior Cloud Engineer to help design, build, and operate platform ... Implement and maintain platform observability and reliability tooling (Prometheus, Grafana, Datadog ...

Sr. AI Engineer

Wilmington, DE ยท On-site

$118K - $156K/yr

Experience with AI Observability in platforms like Datadog and Langsmith. * Bachelor's degree with emphasis in related field or equivalent experience. Qualifications that are not required but are a ...

Senior DevOps Engineer

Camden, NJ ยท On-site

$125K - $145K/yr

Sr DevOps Engineer Who We Are US Cold owns and operates one of the most complex temperature ... Implement monitoring, alerting, logging, and observability solutions to proactively identify and ...

Senior AI Engineer

Glassboro, NJ ยท On-site +1

$101K - $138K/yr

As a Senior Lead Applied AI Engineer, you will design, build, and operate production-grade internal ... observability, and deployment. As a core implementation partner to the AI Product Managers, you ...

next page

Showing results 1-20

Senior Observability Engineer information

See Philadelphia, PA salary details

$60K

$127.7K

$185.2K

How much do senior observability engineer jobs pay per year?

As of Sep 2, 2026, the average yearly pay for senior observability engineer in Philadelphia, PA is $127,707.00, according to ZipRecruiter salary data. Most workers in this role earn between $105,400.00 and $144,800.00 per year, depending on experience, location, and employer.

What is a senior observability engineer?

A Senior Observability Engineer is a seasoned IT professional responsible for designing, implementing, and maintaining systems that monitor and provide insights into the performance, health, and reliability of software applications and infrastructure. They utilize tools for logging, monitoring, tracing, and alerting to ensure that systems are observable and any issues can be quickly detected and resolved. In addition to technical expertise, they often collaborate with development and operations teams to establish best practices, improve incident response, and optimize system performance. Their work is crucial for maintaining uptime, enhancing customer experiences, and supporting the scalability of technology platforms.

How does a senior observability engineer typically collaborate with development and operations teams?

A Senior Observability Engineer works closely with both development and operations teams to ensure robust monitoring, logging, and tracing solutions are in place across all applications and infrastructure. They often participate in architecture discussions to advise on best practices for instrumenting code and systems for observability. By analyzing metrics and alerting patterns, they help teams proactively resolve issues and optimize system performance. This role also involves mentoring engineers on observability tools and fostering a culture of transparency and accountability in incident response.

What are the key skills and qualifications needed to thrive as a senior observability engineer, and why are they important?

To thrive as a Senior Observability Engineer, you need expertise in monitoring, logging, and tracing systems, with a solid background in computer science or a related field. Familiarity with tools like Prometheus, Grafana, ELK stack, and cloud platforms, as well as certifications such as AWS Certified DevOps Engineer, are typically required. Strong problem-solving, collaboration, and communication skills are critical for effectively diagnosing and resolving complex infrastructure issues. These skills ensure reliable system performance, rapid incident response, and continuous improvement of the technology environment.

What is the difference between Senior Observability Engineer vs Site Reliability Engineer?

AspectSenior Observability EngineerSite Reliability Engineer
CredentialsExperience with monitoring tools, scripting, cloud platformsSame as Senior Observability Engineer, often with SRE certifications
Work EnvironmentFocus on monitoring, logging, and tracing systemsFocus on system reliability, automation, and incident response
Industry UsageUsed in tech companies emphasizing system observabilityCommon in large-scale tech and cloud services
Search/Comparison IntentOften compared for monitoring rolesCompared for reliability and system stability roles

While both roles require expertise in cloud platforms and scripting, the Senior Observability Engineer primarily focuses on designing and maintaining monitoring, logging, and tracing systems to ensure system visibility. In contrast, a Site Reliability Engineer emphasizes system reliability, automation, and incident management to maintain service uptime. Both roles are vital in tech environments but serve different core functions related to system health and stability.

How much do senior observability engineers make?

Senior observability engineers typically earn between $110,000 and $160,000 annually, depending on experience, location, and company size. They often work with tools like Prometheus, Grafana, and cloud platforms, and may require advanced knowledge of monitoring, logging, and alerting systems.

What does a senior observability engineer do?

A senior observability engineer designs, implements, and maintains systems to monitor the performance and health of software applications and infrastructure. They utilize tools like Prometheus, Grafana, and ELK stack to analyze metrics, logs, and traces, ensuring system reliability and performance. This role often requires strong scripting skills and knowledge of cloud environments and distributed systems.

What are the most commonly searched types of Observability Engineer jobs in Philadelphia, PA?

The most popular types of Observability Engineer jobs in Philadelphia, PA are:

What are popular job titles related to Senior Observability Engineer jobs in Philadelphia, PA?

For Senior Observability Engineer jobs in Philadelphia, PA, the most frequently searched job titles are:

What job categories do people searching Senior Observability Engineer jobs in Philadelphia, PA look for?

The top searched job categories for Senior Observability Engineer jobs in Philadelphia, PA are:

What cities near Philadelphia, PA are hiring for Senior Observability Engineer jobs?

Cities near Philadelphia, PA with the most Senior Observability Engineer job openings:

Senior Platform Engineer (AI Gateway, Identity & Observability)

TekShapers

Oaks, PA โ€ข On-site

$125K - $165K/yr

Other

This job post hasย expired 1 day ago.ย Applications are no longer accepted.


Job description

Role:                    Senior Platform Engineer (AI Gateway, Identity & Observability)

Location:            Oaks, PA

Type :                    Long Term Contract

Exp. Required: 10+ years

Need candidate who can work onsite from Day 1 (Hybrid basis)

Experience: 10+ years overall, including 5+ years hands-on API platform, identity or observability engineering

RESPONSIBILITIES

The platform is being designed as the engagement progresses, so responsibilities will evolve. Indicative, and not limited to:

  • Deliver the AI gateway that routes model traffic across cloud-hosted models, other providers and on-premises inference, presenting one consistent, provider-agnostic interface to every consuming agent.
  • Implement gateway policy: routing and backend pools, token-based rate limiting and quotas, retries, failover and circuit breaking, caching, and request and response transformation.
  • Harden an existing service for production and run a controlled migration to the long-term gateway, moving consumers across without changing the contract they already integrated against.
  • Build the agent identity and registration path: workload or service identity for agents, registration on deployment, and the promotion gate that governs what reaches production. Full automation is not assumed; part of the work is automating what can be automated and designing a clean, documented path for the steps that require human approval.

Skills Required

  • 5+ years hands-on engineering across API platforms, identity or observability, with production depth in at least one and working competence in the others. All three areas are in scope for this position.
  • 4+ years building and operating API gateways in production: routing, policy authoring, rate limiting and quotas, authentication, transformation and versioning. Azure API Management is strongly preferred; comparable depth in Apigee, Kong or an equivalent enterprise gateway is acceptable.
  • 3+ years enterprise identity and access engineering: OAuth2, OpenID Connect and JWT validation, service principals and workload or managed identity, and secrets management using an enterprise vault.
  • 3+ years hands-on OpenTelemetry or equivalent distributed tracing: instrumentation, collectors, resource attributes, span design, and trace, metric and log pipelines.
  • 3+ years in a site reliability, production engineering or platform operations role, or equivalent hands-on responsibility for a service other teams depend on: service level objectives and error budgets, alerting design, on-call and incident response, and post-incident analysis. This position owns whether the telemetry is trustworthy, not only whether it is being collected.
  • 2+ years working with LLM or AI workloads in production, including model routing across providers, token-based limits and quotas, streaming responses, content filtering, and how inference cost accrues and is attributed.
  • Experience with LLM observability tooling, such as Langfuse, LangSmith, Arize or an equivalent platform, including how agent traces differ from conventional application traces.
  • 3+ years production API operations: failover, circuit breaking, caching, load testing to prove capacity, and incident response for a service other teams depend on.
  • 2+ years integrating third-party or self-hosted services into an enterprise network, including private connectivity or controlled egress, DNS resolution, TLS and certificate management, and working the firewall and security review needed to get each path approved.
  • Working knowledge of cloud infrastructure on at least one major hyperscaler: networking, identity, container platforms and infrastructure-as-code.
  • Programming competence in Python, C# or an equivalent language, sufficient to build and maintain policy extensions, callout services and instrumentation libraries.

PREFERRED

  • Familiarity with Azure AI Foundry observability and Azure Monitor or Application Insights, including agent tracing, continuous evaluation and how a managed observability plane compares with a self-hosted one such as Langfuse. This platform may run one, the other, or both, and the choice is still open.
  • Azure API Management depth, including policy expressions, reusable policy fragments and the AI or LLM policy set.
  • Microsoft Entra ID experience, including managed identity, app registrations and conditional access.
  • Experience with reliability engineering for AI or machine learning systems, including quality regression detection, drift monitoring or evaluation-driven alerting.
  • Familiarity with emerging agent identity models and how agent-to-service authentication differs from conventional service-to-service patterns.
  • Experience implementing chargeback or showback for a shared platform service.
  • Content safety or guardrail services, such as Azure AI Content Safety, or custom filtering and moderation services.
  • Enterprise workflow integration, such as ServiceNow APIs for approval-driven provisioning.
  • Experience in financial services or another regulated industry, where security review governs the pace of change.
  • Familiarity with agent frameworks and how agents consume model endpoints, tools and memory.
Tekshapers is an equal opportunity employer and will consider all applications without regards to race, sex, age, color, religion, national origin, veteran status, disability, sexual orientation, gender identity, genetic information or any characteristic protected by law.