1

Observability Ebpf Jobs (NOW HIRING)

Senior Storage Software Engineer, DGXC Data Services

OR ยท On-site +1

$122K - $161K/yr

Background with Linux kernel observability, eBPF, tracing, or low-overhead telemetry systems. Experience with FUSE, POSIX filesystems, object-store-backed filesystems, or filesystem metadata/indexing.

next page

Showing results 1-20

Observability Ebpf information

See salary details

$16

$60

$86

How much do observability ebpf jobs pay per hour?

As of Sep 9, 2026, the average hourly pay for observability ebpf in the United States is $60.53, according to ZipRecruiter salary data. Most workers in this role earn between $50.72 and $69.47 per hour, depending on experience, location, and employer.

What is an Observability eBPF engineer?

An Observability eBPF Engineer is a specialist who leverages eBPF (extended Berkeley Packet Filter) technology to improve system observability, monitoring, and performance analysis. They design and implement tools that collect deep insights from operating systems and applications without significant overhead. Their work involves writing eBPF programs, integrating them with observability platforms, and ensuring that systems can be monitored securely and efficiently. These engineers often collaborate with SRE, DevOps, and security teams to enable better troubleshooting and incident response. Their expertise is crucial for modern cloud-native environments where traditional monitoring tools may fall short.

How does an Observability eBPF engineer typically collaborate with development and operations teams?

An Observability eBPF engineer works closely with both development and operations teams to instrument applications and infrastructure, ensuring that performance and security metrics are accurately collected in real-time. They often participate in designing monitoring solutions, troubleshooting complex issues, and providing actionable insights that improve system reliability. Regular collaboration involves sharing findings, recommending optimizations, and integrating observability tools into CI/CD pipelines. This role requires effective communication skills to translate technical metrics into meaningful information for broader teams.

What are the key skills and qualifications needed to thrive as an Observability eBPF engineer, and why are they important?

To thrive as an Observability eBPF Engineer, you need a solid background in systems programming, Linux internals, networking, and experience with eBPF technology, often supported by a degree in computer science or a related field. Familiarity with observability tools like Prometheus, Grafana, and distributed tracing systems, as well as proficiency in languages like C, Go, or Rust, is essential. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for diagnosing complex issues and collaborating with cross-functional teams. These skills and qualifications are vital for building reliable, high-performance monitoring solutions and ensuring system observability at scale.

What other helpful pages are available for Observability Ebpf?

Other pages related to Observability Ebpf:

Infographic showing various Observability Ebpf job openings in the United States as of September 2026, with employment types broken down into 96% Full Time, and 4% Contract. Highlights an 74% Physical, 7% Hybrid, and 19% Remote job distribution, with an average salary of $125,908 per year, or $60.5 per hour.

Grafana Observability SME

Manhattan, NY โ€ข On-site

Ruri Software Technologies LLC
Software Developmentย โ€ขย 11 - 50 employees

Full-time

This job post hasย expired 1 day ago.ย Applications are no longer accepted.


Job description

Job Title :Grafana Observability SME
Location :Onsite - Poughkeepsie, NY
Job Description:
1. Production expertise across the full Grafana stack: Mimir, Loki, Tempo, Alloy, Beyla, Grafana Application Observability, Unified Alerting.
2. Strong PromQL, LogQL, and TraceQL authoring skills; able to write recording rules and SLO queries from scratch.
3. OpenTelemetry practitioner - OTLP, collectors, SDK/agent instrumentation for at least three of Java, .NET, Go, Python, Node.js.
4. eBPF-based auto-instrumentation experience with Beyla (or equivalent - Pixie, Cilium Tetragon) in a production context.
5. Experience integrating Grafana alerts into ServiceNow Event Management (native inbound integration, not webhook-only patterns); familiarity with ServiceNow ITOM, AIOps event correlation, and CMDB CI attachment.
6. Multi-environment hosting fluency - on-prem, AWS, Azure - and Linux/Windows host agent deployment at scale.
7. Dashboard-as-code and GitOps patterns (Grafana provisioning, Terraform provider, or Grizzly).
8. Excellent written communication - solution architecture documents, runbooks, and stakeholder-facing status reporting.
Role Summary
Own the end-to-end technical design, build, and operationalization of the Grafana Cloud observability platform for a 50-application estate spanning Java, .NET, Go, Python, and Node.js workloads hosted across on-premises data centres, AWS, and Azure. The SME serves as the senior technical authority across all eight in-scope Grafana Cloud modules and is accountable for instrumentation strategy, alerting design, dashboarding standards, and integration into ServiceNow ITOM via native Event Management. Scope is application-level observability only - server and network health remain on SolarWinds, and URL/synthetic monitoring remains on Uptrends.
Key Responsibilities
โ€ข Platform architecture and configuration across all eight in-scope Grafana Cloud modules: Grafana 12 (visualization), Mimir (metrics, 13-month retention), Loki (logs), Tempo (distributed tracing via OTLP), Alloy (telemetry collection agent), Beyla (eBPF zero-code auto-instrumentation), Application Observability (OTel-native APM), and Unified Alerting.
โ€ข Tenancy and access design - organizations, folders, teams, role-based access control, dashboard variables, template links, and annotations.
โ€ข Application instrumentation strategy by technology stack: Beyla eBPF as the default zero-code path for Simple and Medium apps; OpenTelemetry SDKs/agents (Java, .NET, Go, Python, Node.js) for Complex apps requiring deeper traces and custom metrics; JMX Exporter, prometheus_client, and runtime-specific exporters where stack-appropriate.
โ€ข Log pipeline engineering via Alloy - structured JSON, Log4j/Logback, Serilog, NLog, Windows Event Log, Winston, Pino, loguru - with parsing rules tuned per stack and LogQL-based dashboards and alerts.
โ€ข Alerting design - PromQL/LogQL/TraceQL rules, severity taxonomy, grouping, routing, and notification policies. Build a low-noise, actionable alert feed; tune thresholds iteratively with application owners.
โ€ข Single Pane of Glass - design and deliver a tiered SPoG that surfaces Grafana application telemetry alongside contextual links to SolarWinds and Uptrends.
โ€ข Business Dashboards and Reporting - partner with the Dashboard Lead to define KPI taxonomy and ensure dashboard-as-code patterns and version control.
โ€ข ServiceNow ITOM integration - co-own the design and review of Grafana โ†’ ServiceNow Event Management (native inbound integration) flow: event allow-list governance ("deny by default"), enrichment, deduplication, AIOps correlation, automated incident creation with severity mapping and assignment group rules, CMDB CI attachment, and ServiceNow-as-master incident state.
โ€ข Quality assurance authority across all technical deliverables - solution architecture document, instrumentation runbooks, dashboard and alert library, integration test results.
โ€ข Phased delivery execution - Mobilise & Discover โ†’ Application Foundation (ML1) โ†’ Onboarding of 40 Simple apps (ML2) โ†’ Medium/Complex apps + ITOM Integration (ML2โ†’3) โ†’ SPoG, Dashboards & Reporting (ML3โ†’4) โ†’ Stabilisation, KT, and post-deployment support (ML4).
โ€ข Knowledge transfer - produce platform operating procedures and conduct structured handover to the client's run team.
Required Skills & Experience
โ€ข 7+ years in observability/monitoring engineering with deep, recent hands-on Grafana Cloud experience (not just OSS Grafana).
โ€ข Production expertise across the full Grafana stack: Mimir, Loki, Tempo, Alloy, Beyla, Grafana Application Observability, Unified Alerting.
โ€ข Strong PromQL, LogQL, and TraceQL authoring skills; able to write recording rules and SLO queries from scratch.
โ€ข OpenTelemetry practitioner - OTLP, collectors, SDK/agent instrumentation for at least three of Java, .NET, Go, Python, Node.js.
โ€ข eBPF-based auto-instrumentation experience with Beyla (or equivalent - Pixie, Cilium Tetragon) in a production context.
โ€ข Experience integrating Grafana alerts into ServiceNow Event Management (native inbound integration, not webhook-only patterns); familiarity with ServiceNow ITOM, AIOps event correlation, and CMDB CI attachment.
โ€ข Multi-environment hosting fluency - on-prem, AWS, Azure - and Linux/Windows host agent deployment at scale.
โ€ข Dashboard-as-code and GitOps patterns (Grafana provisioning, Terraform provider, or Grizzly).
โ€ข Excellent written communication - solution architecture documents, runbooks, and stakeholder-facing status reporting.
Nice to Have
โ€ข Grafana Certified Professional or equivalent vendor certification.
โ€ข Prior experience in a regulated utility, energy, or critical-infrastructure environment.
โ€ข Familiarity with SolarWinds and Uptrends (sufficient to design clean boundaries with retained tooling, not to administer them).
โ€ข Experience with ServiceNow CSDM and Service Mapping governance.
โ€ข Exposure to FinOps for observability - cardinality control, log volume management, retention tuning in Mimir/Loki.
Out of Scope for This Role
โ€ข Server health and network monitoring (owned by SolarWinds).
โ€ข URL/synthetic endpoint monitoring (owned by Uptrends).
โ€ข ServiceNow ITSM workflow ownership - incident lifecycle remains with the client's ITSM/ITOM team; this role designs the integration, not the downstream process.
Years of Experience: 12.00 Years of Experience
Regards
Surya
surya@rurisoft.com