1

Observability Sre Jobs (NOW HIRING)

Site Reliability Engineer (SRE)

Plano, TX · On-site

$54.50 - $72.50/hr

Site Reliability Engineer (SRE) Location: Richmond, VA or Plano, TX Work Model: Hybrid - 3 days ... Observability tools (Prometheus, Grafana, Datadog, CloudWatch) * Experience working in highly ...

Site Reliability Engineer

Santa Clara, CA · On-site

$230K - $250K/yr

Build and maintain observability infrastructure - logging, metrics, tracing, and alerting - so the ... Help define and build the SRE team as the company scales - this is a foundational hire with a path ...

Site Reliability Engineer

Beaverton, OR · Hybrid

$59.25 - $78.75/hr

Overview As our Site Reliability Engineer, you'll help drive Concora Credit's Mission to enable ... Contribute to cloud reliability through automation, observability, incident reduction, capacity ...

Build and maintain observability infrastructure - logging, metrics, tracing, and alerting - so the ... Help define and build the SRE team as the company scales - this is a foundational hire with a path ...

Site Reliability Engineer

Beaverton, OR · Hybrid

$59.25 - $78.75/hr

Overview As our Site Reliability Engineer, you'll help drive Concora Credit's Mission to enable ... Contribute to cloud reliability through automation, observability, incident reduction, capacity ...

Site Reliability Engineer

Beaverton, OR · Hybrid

$59.25 - $78.75/hr

As our Site Reliability Engineer, you'll help drive Concora Credit's Mission to enable customers to ... Automation, Observability, and Continuous Improvement: • Improve operational efficiency through ...

Site Reliability Engineer

Plano, TX · On-site

$54.50 - $72.50/hr

Site Reliability Engineering (SRE): • Ensure high availability, scalability, and reliability of distributed systems. • Implement observability, logging, and monitoring using tools like Prometheus ...

Observability adoption (OpenTelemetry, Dynatrace) * Reliability engineering practices * Platform standardization and automation The ideal candidate combines software engineering expertise with SRE ...

This role focuses on automation, observability, and incident response while upholding strict Service Level Objectives (SLOs). The SRE will help build resilient systems that scale, automate manual ...

Site Reliability Engineer, Observability

Chicago, IL · On-site

$58.75 - $78/hr

Core SRE Experience * 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with a strong focus on observability and production operations. * Proven ability to deliver hands-on ...

Site Reliability Engineer, Observability

New York, NY · On-site

$62.25 - $82.75/hr

Core SRE Experience * 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with a strong focus on observability and production operations. * Proven ability to deliver hands-on ...

next page

Showing results 1-20

Observability Sre information

See salary details

$10

$63

$91

How much do observability sre jobs pay per hour?

As of Jul 24, 2026, the average hourly pay for observability sre in the United States is $63.74, according to ZipRecruiter salary data. Most workers in this role earn between $54.81 and $72.84 per hour, depending on experience, location, and employer.

What is the difference between Observability Sre vs Site Reliability Engineer?

AspectObservability SreSite Reliability Engineer
Primary FocusMonitoring, logging, and tracing to ensure system observabilitySystem reliability, automation, and infrastructure management
Skills & CertificationsMonitoring tools, scripting, cloud platforms, observability frameworksLinux, scripting, cloud services, automation tools
Work EnvironmentCollaborates with SRE, DevOps, and development teams on observability practicesBuilds and maintains scalable, reliable systems in production

While both roles focus on system stability, Observability Sre specializes in monitoring and diagnostics, whereas Site Reliability Engineers focus on overall system reliability and automation. They often work together to ensure robust, observable, and reliable systems.

What is an Observability SRE?

An Observability SRE (Site Reliability Engineer) is a specialist focused on ensuring that systems and applications are transparent, measurable, and reliable. Their main responsibility is to implement and maintain tools for monitoring, logging, and tracing, providing insights into system performance and health. Observability SREs help teams quickly detect, diagnose, and resolve issues by making system behavior visible and understandable. They play a critical role in uptime, incident response, and performance optimization, bridging the gap between software development and IT operations.

What are some typical challenges faced by Observability SREs when implementing monitoring solutions across diverse systems?

Observability SREs often encounter challenges when integrating monitoring tools across varied technology stacks and legacy systems. Ensuring consistent data collection, standardizing metrics, and maintaining visibility in complex, distributed environments can be difficult. Collaborating with development and operations teams to define meaningful alerts and dashboards requires strong communication and a deep understanding of both infrastructure and application behaviors. Staying up-to-date with evolving tools and best practices is also essential to address emerging observability needs.

What are the key skills and qualifications needed to thrive as an Observability SRE, and why are they important?

To thrive as an Observability SRE, you need a solid background in systems engineering, monitoring best practices, and expertise in observability concepts, often supported by a degree in computer science or related fields. Familiarity with tools like Prometheus, Grafana, ELK stack, and cloud monitoring platforms, as well as scripting languages such as Python or Bash, is typically required. Strong problem-solving, collaboration, and communication skills help SREs respond to incidents and work across teams effectively. These skills ensure system reliability, rapid issue detection, and continuous service improvement in complex technical environments.
More about Observability Sre jobs
What cities are hiring for Observability Sre jobs? Cities with the most Observability Sre job openings:
What states have the most Observability Sre jobs? States with the most job openings for Observability Sre jobs include:
Infographic showing various Observability Sre job openings in the United States as of July 2026, with employment types broken down into 97% Full Time, and 3% Contract. Highlights an 76% Physical, 7% Hybrid, and 17% Remote job distribution, with an average salary of $132,583 per year, or $63.7 per hour.
Site Reliability Engineer (SRE)

Site Reliability Engineer (SRE)

IT America Inc

Plano, TX • On-site

$54.50 - $72.50/hr

Contractor

Posted 19 days ago


Job description

Position: Site Reliability Engineer (SRE)

Location: Richmond, VA or Plano, TX

Work Model: Hybrid – 3 days onsite per week

Duration: Long term contract

Job Summary:

We are seeking an experienced Site Reliability Engineer (SRE) to support cloud-native platforms and production systems for a large enterprise environment. This role will focus on ensuring high availability, reliability, performance, and scalability of mission-critical applications running on AWS.

Strong Preference: Former Capital One engineers. Candidates must be able to provide verifiable Capital One credentials and be eligible for rehire.

Key Responsibilities:

  • Design, build, and maintain highly reliable, scalable, and resilient systems in AWS
  • Monitor system health, performance, and availability using SRE best practices
  • Implement automation to reduce manual operational work
  • Troubleshoot production incidents and perform root cause analysis (RCA)
  • Develop and maintain scripts and tools to improve system reliability and efficiency
  • Partner with application development, platform, and infrastructure teams
  • Support on-call rotations and incident response as required
  • Enforce operational excellence, security, and compliance standards

Required Skills & Qualifications:

  • Former Capital One experience – HIGHLY preferred
  • Must provide credentials for rehire eligibility verification
  • Strong hands-on experience with AWS (EC2, EKS, Lambda, CloudWatch, IAM, etc.)
  • Python scripting experience strongly preferred
  • Bash or Shell scripting experience will also be considered
  • Experience with Linux-based systems and troubleshooting
  • Understanding of SRE concepts: SLIs, SLOs, error budgets, monitoring, and alerting
  • Experience supporting production environments at scale

Preferred Qualifications:

  • Experience with CI/CD pipelines
  • Infrastructure as Code (Terraform, CloudFormation)
  • Containerization and orchestration (Docker, Kubernetes)
  • Observability tools (Prometheus, Grafana, Datadog, CloudWatch)
  • Experience working in highly regulated enterprise environments