1

Observability Sre Jobs (NOW HIRING)

Staff SRE - Observability

Chicago, IL

$58.75 - $78/hr

Build CI/CD pipelines with embedded observability and automated testing Site Reliability Engineering (SRE) * Establish and maintain Service Level Indicators (SLIs), Objectives (SLOs), and Agreements ...

Staff SRE - Observability

Chicago, IL · On-site

$58.75 - $78/hr

Build CI/CD pipelines with embedded observability and automated testing Site Reliability Engineering (SRE) * Establish and maintain Service Level Indicators (SLIs), Objectives (SLOs), and Agreements ...

Site Reliability Engineer(SRE)

Dallas, TX · On-site

$56.75 - $75.25/hr

This role focuses on automation, observability, and proactive incident prevention to ensure high ... SRE / Production Support / IT Operations * Strong scripting/automation expertise ( Python ...

SITE RELIABILITY ENGINEER

Camden, NJ · On-site

$130K - $150K/yr

Site Reliability Engineer (SRE) Engineer Reliability into the Systems That Move the Nation's Food ... Observability across the full stack, correlating cloud services, APIs, and on-premise facility ...

Site Reliability Engineer

Beaverton, OR · On-site

$59.25 - $78.75/hr

Overview As our Site Reliability Engineer, you'll help drive Concora Credit's Mission to enable ... Automation, Observability, and Continuous Improvement: • Improve operational efficiency through ...

next page

Showing results 1-20

Observability Sre information

See salary details

$10

$63

$91

How much do observability sre jobs pay per hour?

As of Jul 23, 2026, the average hourly pay for observability sre in the United States is $63.74, according to ZipRecruiter salary data. Most workers in this role earn between $54.81 and $72.84 per hour, depending on experience, location, and employer.

What is the difference between Observability Sre vs Site Reliability Engineer?

AspectObservability SreSite Reliability Engineer
Primary FocusMonitoring, logging, and tracing to ensure system observabilitySystem reliability, automation, and infrastructure management
Skills & CertificationsMonitoring tools, scripting, cloud platforms, observability frameworksLinux, scripting, cloud services, automation tools
Work EnvironmentCollaborates with SRE, DevOps, and development teams on observability practicesBuilds and maintains scalable, reliable systems in production

While both roles focus on system stability, Observability Sre specializes in monitoring and diagnostics, whereas Site Reliability Engineers focus on overall system reliability and automation. They often work together to ensure robust, observable, and reliable systems.

What is an Observability SRE?

An Observability SRE (Site Reliability Engineer) is a specialist focused on ensuring that systems and applications are transparent, measurable, and reliable. Their main responsibility is to implement and maintain tools for monitoring, logging, and tracing, providing insights into system performance and health. Observability SREs help teams quickly detect, diagnose, and resolve issues by making system behavior visible and understandable. They play a critical role in uptime, incident response, and performance optimization, bridging the gap between software development and IT operations.

What are some typical challenges faced by Observability SREs when implementing monitoring solutions across diverse systems?

Observability SREs often encounter challenges when integrating monitoring tools across varied technology stacks and legacy systems. Ensuring consistent data collection, standardizing metrics, and maintaining visibility in complex, distributed environments can be difficult. Collaborating with development and operations teams to define meaningful alerts and dashboards requires strong communication and a deep understanding of both infrastructure and application behaviors. Staying up-to-date with evolving tools and best practices is also essential to address emerging observability needs.

What are the key skills and qualifications needed to thrive as an Observability SRE, and why are they important?

To thrive as an Observability SRE, you need a solid background in systems engineering, monitoring best practices, and expertise in observability concepts, often supported by a degree in computer science or related fields. Familiarity with tools like Prometheus, Grafana, ELK stack, and cloud monitoring platforms, as well as scripting languages such as Python or Bash, is typically required. Strong problem-solving, collaboration, and communication skills help SREs respond to incidents and work across teams effectively. These skills ensure system reliability, rapid issue detection, and continuous service improvement in complex technical environments.
More about Observability Sre jobs
What cities are hiring for Observability Sre jobs? Cities with the most Observability Sre job openings:
What states have the most Observability Sre jobs? States with the most job openings for Observability Sre jobs include:
Infographic showing various Observability Sre job openings in the United States as of July 2026, with employment types broken down into 97% Full Time, and 3% Contract. Highlights an 76% Physical, 7% Hybrid, and 17% Remote job distribution, with an average salary of $132,583 per year, or $63.7 per hour.

Staff SRE - Observability

Focused

Chicago, IL

$58.75 - $78/hr

Other

Posted yesterday


Job description

Who we are:

At Focused, we move quickly to deliver quality software that achieves client outcomes and meets their customer's needs. We strategically partner with our clients to leverage our expertise in design and software, while our clients bring their own domain expertise. We work with a variety of clients from different industries, collaborating as we get new products to market, modernizing legacy systems, or helping teams learn the skills they need to be successful.   

Our values:

  • Listen first  We are experts in product practices but life long learners in the domain of our customers. We research, collaborate, and understand. 
  • Learn why  We ask questions and talk to users to understand problem spaces, objectives, and goals, which allows us to deeply invest and drive towards the outcomes of our clients. 
  • Love your craft  We love diving into a variety of domains and solving problems.  We take pride in delivering value, in communicating progress, and guiding our clients to success.

We are seeking an experienced Staff Observability Consultant with deep expertise in OpenTelemetry and strong Platform Engineering capabilities to help organizations implement, optimize, and scale their observability infrastructure. This role requires a seasoned consultant who can design comprehensive telemetry strategies, implement distributed tracing solutions, establish robust monitoring practices, and interface closely with clients on the observability journey.

Key Responsibilities:

OpenTelemetry & Observability

  • Design and implement end-to-end OpenTelemetry solutions across diverse technology stacks
  • Configure and deploy OpenTelemetry Collectors for efficient data collection, processing, sampling, and routing
  • Establish telemetry pipelines for metrics, traces, and logs across microservices architectures
  • Optimize collector configurations for performance, reliability, and cost-effectiveness

Platform Engineering & Infrastructure

  • Augment existing infrastructure with with integrated observability solutions
  • Implement Infrastructure as Code (IaC) solutions using Terraform, Pulumi, CloudFormation, etc.
  • Architect and manage Kubernetes clusters with comprehensive monitoring and logging
  • Build CI/CD pipelines with embedded observability and automated testing

Site Reliability Engineering (SRE)

  • Establish and maintain Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs)
  • Implement error budgets, toil reduction strategies, and capacity planning
  • Support incident response procedures and post-mortem processes

Cloud & DevOps Engineering

  • Deploy and manage observability infrastructure across AWS, GCP, and Azure
  • Establish security, compliance, and governance frameworks for telemetry data
  • Experience automating Agent Evaluations in CI/CD pipelines and observability backends.

Required Qualifications:

Core Observability & OpenTelemetry

  • 3-7 years of experience in observability, monitoring, and distributed systems
  • Deep hands-on experience with OpenTelemetry ecosystem, including SDKs, APIs, and specifications
  • Proficiency with OpenTelemetry Collector configuration, processors, exporters, and receivers
  • Strong understanding of telemetry data models, semantic conventions, and instrumentation best practices

Platform Engineering & DevOps

  • 5+ years of Platform Engineering or DevOps experience with focus on site reliability, observability, and incident response
  • Proficiency with Infrastructure as Code tools (Terraform, Pulumi, CloudFormation, CDK)
  • Strong experience with CI/CD platforms (GitHub Actions, GitLab CI, Jenkins, ArgoCD)

Cloud & Infrastructure

  • Hands-on experience with major cloud providers (AWS, GCP, Azure) and their observability services
  • Experience with container technologies (Docker, Podman) and container registries
  • Knowledge of networking, security, load balancing, and distributed systems concepts

Site Reliability Engineering

  • Experience implementing SRE practices including error budgets and toil metrics
  • Proficiency in incident management, on-call procedures, and post-mortem culture
  • Experience with capacity planning, performance optimization, and scalability design

Programming & Automation

  • Proficiency in multiple programming languages preferred (Go, Python, Java, Node.js, Rust)
  • Strong scripting and automation skills (Bash, Python, PowerShell)
  • Understanding of software engineering best practices and testing methodologies

Preferred Qualifications (Exceptional Candidates)

AI & Agentic Frameworks

  • Understanding of Large Language Models (LLMs) and their application in DevOps
  • Knowledge of vector databases, embeddings, and retrieval-augmented generation (RAG)
  • Experience with AI/ML model deployment and monitoring in production environments

Leadership & Communication

  • Strong technical writing and documentation skills
  • Ability to present complex technical concepts to diverse stakeholders
  • A passion for knowledge sharing

Key Competencies

  • Systems thinking and ability to design holistic observability solutions
  • Strong analytical and troubleshooting skills for complex distributed systems
  • Curiosity about emerging technologies, particularly AI applications in operations
  • Adaptability to rapidly evolving cloud-native and observability technologies
  • Collaborative mindset with focus on enabling developer productivity and system reliability

What Sets Exceptional Candidates Apart:

  • Experience with Honeycomb
  • Contributions to open-source observability or AI framework projects
  • Track record of implementing platform engineering solutions that significantly improved developer experience
  • Experience scaling observability infrastructure to handle high event volume

What to know before you apply: 

  • This role will require being in the Chicago office three days per week and up to 20% travel within the United States.
  • Focused is unable to sponsor or take over sponsorship of the employment Visa process at this time.
  • The Chicago base salary range for this role is $160,000 - $200,000.