1

Cloud Observability Manager Jobs (NOW HIRING)

Sr. AWS Cloud Engineer

Irvine, CA ยท On-site

$59.25 - $79/hr

... state management. โ€ข Familiarity with Java/.NET apps, TIBCO ESB, and infra dependencies during ... observability. โ€ข Ability to operate within cloud governance, cost control, and compliance ...

Cloud Engineer

Indianapolis, IN ยท On-site

$187K - $218K/yr

... management, while ensuring alignment with regulatory standards (e.g., GxP, HIPAA). * Monitoring & Optimization: Oversee cloud observability using tools like Prometheus and Grafana, and apply FinOps ...

Cloud Engineer

$57 - $76.25/hr

Familiarity with cloud monitoring and observability tools such as CloudWatch, Prometheus, or Datadog. * 2+ years of experience working within ITSM and change management processes. * Strong verbal and ...

Cloud Engineer

$57 - $76.25/hr

Familiarity with cloud monitoring and observability tools such as CloudWatch, Prometheus, or Datadog. * 2+ years of experience working within ITSM and change management processes. * Strong verbal and ...

Sr. Cloud Architect

$66.50 - $84.75/hr

Develops and guides cloud observability strategies and optimize AWS environments for performance ... Partners with product management, engineering, cloud operations, architecture, and security teams ...

Cloud Observability & Tooling * Implement, administer, and optimize Flexera One and native cloud cost-management platforms. * Develop detailed dashboards that accurately allocate shared ...

Cloud Observability & Tooling * Implement, administer, and optimize Flexera One and native cloud cost-management platforms. * Develop detailed dashboards that accurately allocate shared ...

Cloud Engineer

Lake Mary, FL ยท Remote

$57 - $76.25/hr

Cloud Infrastructure & Operations โ— Administer the AWS Organization, managing accounts ... observability, including CloudWatch metrics, logs, and alarms, and CloudTrail, and respond to ...

Cloud Engineer

Lake Mary, FL ยท On-site +1

$48.75 - $65.25/hr

Cloud Infrastructure & Operations โ€ข Administer the AWS Organization, managing accounts ... observability, including CloudWatch metrics, logs, and alarms, and CloudTrail, and respond to ...

Cloud Engineer III

$57 - $76.25/hr

Proven experience with cloud cost management, observability, and performance optimization * Demonstrated technical leadership and mentoring skills * Extensive governance experience (Azure Policy ...

Showing results 41-60

Cloud Observability Manager information

See salary details

$23K

$61.4K

$102.5K

How much do cloud observability manager jobs pay per year?

As of Sep 14, 2026, the average yearly pay for cloud observability manager in the United States is $61,351.00, according to ZipRecruiter salary data. Most workers in this role earn between $44,000.00 and $69,000.00 per year, depending on experience, location, and employer.

What are popular job titles related to Cloud Observability Manager jobs?

For Cloud Observability Manager jobs, the most frequently searched job titles are:

Grafana Observability SME

Manhattan, NY โ€ข On-site

Ruri Software Technologies LLC
Software Developmentย โ€ขย 11 - 50 employees

Full-time

Re-posted 27 days ago


Job description

Job Title :Grafana Observability SME
Location :Onsite - Poughkeepsie, NY
Job Description:
1. Production expertise across the full Grafana stack: Mimir, Loki, Tempo, Alloy, Beyla, Grafana Application Observability, Unified Alerting.
2. Strong PromQL, LogQL, and TraceQL authoring skills; able to write recording rules and SLO queries from scratch.
3. OpenTelemetry practitioner - OTLP, collectors, SDK/agent instrumentation for at least three of Java, .NET, Go, Python, Node.js.
4. eBPF-based auto-instrumentation experience with Beyla (or equivalent - Pixie, Cilium Tetragon) in a production context.
5. Experience integrating Grafana alerts into ServiceNow Event Management (native inbound integration, not webhook-only patterns); familiarity with ServiceNow ITOM, AIOps event correlation, and CMDB CI attachment.
6. Multi-environment hosting fluency - on-prem, AWS, Azure - and Linux/Windows host agent deployment at scale.
7. Dashboard-as-code and GitOps patterns (Grafana provisioning, Terraform provider, or Grizzly).
8. Excellent written communication - solution architecture documents, runbooks, and stakeholder-facing status reporting.
Role Summary
Own the end-to-end technical design, build, and operationalization of the Grafana Cloud observability platform for a 50-application estate spanning Java, .NET, Go, Python, and Node.js workloads hosted across on-premises data centres, AWS, and Azure. The SME serves as the senior technical authority across all eight in-scope Grafana Cloud modules and is accountable for instrumentation strategy, alerting design, dashboarding standards, and integration into ServiceNow ITOM via native Event Management. Scope is application-level observability only - server and network health remain on SolarWinds, and URL/synthetic monitoring remains on Uptrends.
Key Responsibilities
โ€ข Platform architecture and configuration across all eight in-scope Grafana Cloud modules: Grafana 12 (visualization), Mimir (metrics, 13-month retention), Loki (logs), Tempo (distributed tracing via OTLP), Alloy (telemetry collection agent), Beyla (eBPF zero-code auto-instrumentation), Application Observability (OTel-native APM), and Unified Alerting.
โ€ข Tenancy and access design - organizations, folders, teams, role-based access control, dashboard variables, template links, and annotations.
โ€ข Application instrumentation strategy by technology stack: Beyla eBPF as the default zero-code path for Simple and Medium apps; OpenTelemetry SDKs/agents (Java, .NET, Go, Python, Node.js) for Complex apps requiring deeper traces and custom metrics; JMX Exporter, prometheus_client, and runtime-specific exporters where stack-appropriate.
โ€ข Log pipeline engineering via Alloy - structured JSON, Log4j/Logback, Serilog, NLog, Windows Event Log, Winston, Pino, loguru - with parsing rules tuned per stack and LogQL-based dashboards and alerts.
โ€ข Alerting design - PromQL/LogQL/TraceQL rules, severity taxonomy, grouping, routing, and notification policies. Build a low-noise, actionable alert feed; tune thresholds iteratively with application owners.
โ€ข Single Pane of Glass - design and deliver a tiered SPoG that surfaces Grafana application telemetry alongside contextual links to SolarWinds and Uptrends.
โ€ข Business Dashboards and Reporting - partner with the Dashboard Lead to define KPI taxonomy and ensure dashboard-as-code patterns and version control.
โ€ข ServiceNow ITOM integration - co-own the design and review of Grafana โ†’ ServiceNow Event Management (native inbound integration) flow: event allow-list governance ("deny by default"), enrichment, deduplication, AIOps correlation, automated incident creation with severity mapping and assignment group rules, CMDB CI attachment, and ServiceNow-as-master incident state.
โ€ข Quality assurance authority across all technical deliverables - solution architecture document, instrumentation runbooks, dashboard and alert library, integration test results.
โ€ข Phased delivery execution - Mobilise & Discover โ†’ Application Foundation (ML1) โ†’ Onboarding of 40 Simple apps (ML2) โ†’ Medium/Complex apps + ITOM Integration (ML2โ†’3) โ†’ SPoG, Dashboards & Reporting (ML3โ†’4) โ†’ Stabilisation, KT, and post-deployment support (ML4).
โ€ข Knowledge transfer - produce platform operating procedures and conduct structured handover to the client's run team.
Required Skills & Experience
โ€ข 7+ years in observability/monitoring engineering with deep, recent hands-on Grafana Cloud experience (not just OSS Grafana).
โ€ข Production expertise across the full Grafana stack: Mimir, Loki, Tempo, Alloy, Beyla, Grafana Application Observability, Unified Alerting.
โ€ข Strong PromQL, LogQL, and TraceQL authoring skills; able to write recording rules and SLO queries from scratch.
โ€ข OpenTelemetry practitioner - OTLP, collectors, SDK/agent instrumentation for at least three of Java, .NET, Go, Python, Node.js.
โ€ข eBPF-based auto-instrumentation experience with Beyla (or equivalent - Pixie, Cilium Tetragon) in a production context.
โ€ข Experience integrating Grafana alerts into ServiceNow Event Management (native inbound integration, not webhook-only patterns); familiarity with ServiceNow ITOM, AIOps event correlation, and CMDB CI attachment.
โ€ข Multi-environment hosting fluency - on-prem, AWS, Azure - and Linux/Windows host agent deployment at scale.
โ€ข Dashboard-as-code and GitOps patterns (Grafana provisioning, Terraform provider, or Grizzly).
โ€ข Excellent written communication - solution architecture documents, runbooks, and stakeholder-facing status reporting.
Nice to Have
โ€ข Grafana Certified Professional or equivalent vendor certification.
โ€ข Prior experience in a regulated utility, energy, or critical-infrastructure environment.
โ€ข Familiarity with SolarWinds and Uptrends (sufficient to design clean boundaries with retained tooling, not to administer them).
โ€ข Experience with ServiceNow CSDM and Service Mapping governance.
โ€ข Exposure to FinOps for observability - cardinality control, log volume management, retention tuning in Mimir/Loki.
Out of Scope for This Role
โ€ข Server health and network monitoring (owned by SolarWinds).
โ€ข URL/synthetic endpoint monitoring (owned by Uptrends).
โ€ข ServiceNow ITSM workflow ownership - incident lifecycle remains with the client's ITSM/ITOM team; this role designs the integration, not the downstream process.
Years of Experience: 12.00 Years of Experience
Regards
Surya
surya@rurisoft.com