1

Observability Kubernetes Jobs (NOW HIRING)

Kubernetes Engineer

Phoenix, AZ · On-site

$56.50 - $75.25/hr

Implement observability stacks using tools such as Prometheus, Grafana, Alertmanager, Splunk, ELK ... Apply expertise in Kubernetes networking concepts, including Ingress, Services, CNI plugins, and ...

Experience with CI/CD, observability, Kubernetes, cloud platforms, and enterprise security frameworks. Banking or financial services experience. Expected Outcomes Automated guardrail and adversarial ...

The role requires a strong understanding of observability tools and practices, with a focus on Prometheus, Grafana, Gardener Kubernetes, and Splunk. Experience with Dynatrace is a plus. Skills: • ...

Observability Engineer

Phoenix, AZ · On-site

$52.50 - $71.75/hr

Observability Engineer (Dynatrace, Splunk & OpenSearch) Location: Phoenix, AZ (Onsite) Long Term ... Implement monitoring for cloud-native applications, containers, Kubernetes, and microservices.

Senior Observability Engineer

Natick, MA · On-site

$122K - $189K/yr

Build scalable, multi-tenant observability solutions for Kubernetes clusters running microservices at scale. * Implement SLOs, SLIs, and error budgets-integrating observability into SRE practices.

Kubernetes Architect

Herndon, VA · On-site

$60 - $70/hr

Kubernetes Architect Location-Type: Hybrid (Herndon, VA, Austin, TX, or Newtown Square, PA ... Implement secure monitoring and observability practices across cloud environments * Present ...

next page

Showing results 1-20

Observability Kubernetes information

What cities are hiring for Observability Kubernetes jobs?

Cities with the most Observability Kubernetes job openings:

What states have the most Observability Kubernetes jobs?

States with the most job openings for Observability Kubernetes jobs include:

What job categories do people searching Observability Kubernetes jobs look for?

The top searched job categories for Observability Kubernetes jobs are:

Only W2: SRE / Production Reliability Engineer, Woonsocket, RI

Tror

Woonsocket, RI • On-site

$54.50 - $72.50/hr

Other

This job post has expired today. Applications are no longer accepted.


Job description

Role: Senior SRE / Production Reliability Engineer
Experience: 8+ Years
Location: Woonsocket, RI
 
Job Summary
We are looking for a Senior SRE / Production Reliability Engineer to improve the reliability, performance, and availability of critical production systems.
The ideal candidate should have strong experience in SRE/DevOps, Incident Management, Observability, Kubernetes, Google Cloud Platform, and production monitoring, along with hands-on experience with time-series anomaly detection.
 
What You’ll Do
  • Own reliability and performance of critical production applications.
  • Act as an Incident Commander (IC) during P1/P2 production incidents.
  • Lead incident response, root-cause analysis, postmortems, and reliability improvements.
  • Define and manage SLIs, SLOs, and error budgets.
  • Build and improve monitoring, alerting, and observability solutions.
  • Tune and validate time-series anomaly detection models for production monitoring.
  • Develop automation and operational tools using Python, Java, and React.
  • Troubleshoot Kubernetes, cloud, batch processing, and data pipeline issues.
  • Work with engineering and operations teams to improve system reliability and reduce manual work.
  • Support large-scale deployments and manage production risks such as configuration drift and blast radius.
 
Must-Have Skills
  • 8+ years of experience in SRE, DevOps, Platform Engineering, or Production Engineering.
  • Hands-on experience as an Incident Commander for P1/P2 incidents.
  • Strong experience with time-series anomaly detection models in production observability — mandatory.
  • Strong production-level programming skills in:
    • Python
    • Java
    • React
  • Strong experience with SLI, SLO, and error budgets.
  • Hands-on observability experience with:
    • Prometheus
    • Grafana
    • OpenTelemetry
    • At least 2 log platforms such as Loki, Splunk, or Elasticsearch
  • Strong Google Cloud Platform experience.
  • Strong Kubernetes operational experience.
  • Experience with Rancher K3s.
  • Experience troubleshooting Apache Airflow and Tidal workflows/batch jobs.
  • Experience with production-scale distributed systems and on-call support.
 
Preferred Skills
  • Production Readiness Reviews / service launch experience.
  • BigQuery and PostgreSQL.
  • Chaos Engineering / fault injection.
  • TIC/Technical Incident Commander certification.
  • Healthcare, pharmacy, retail, or other high-availability environments.
  • LLM/GenAI for incident management, alert summarization, or runbook recommendations.
  • Kafka.
  • Istio / Envoy.
  • Terraform / Ansible.
 
Important: The Incident Commander experience and production time-series anomaly detection should be treated as hard requirements. A candidate who only has general monitoring/observability experience but has never worked with anomaly detection models would not be a strong fit.