1

Observability Site Reliability Engineer Jobs in Houston, TX

Site Reliability Engineer

Houston, TX ยท On-site

$55.25 - $73.50/hr

As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident ... Observability & Insights - Set up and maintain full observability stacks (logging, metrics, tracing ...

Site Reliability Engineer

Houston, TX ยท On-site

$55.25 - $73.50/hr

As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident ... Observability & Insights - Set up and maintain full observability stacks (logging, metrics, tracing ...

Site Reliability Engineer

Houston, TX ยท On-site

$55.25 - $73.50/hr

As a Site Reliability Engineer, you will be responsible for: Operational Excellence & Incident ... Observability & Insights - Set up and maintain full observability stacks (logging, metrics, tracing ...

Senior Site Reliability Engineer NEX

Houston, TX ยท On-site

$54.50 - $72.25/hr

Observability and Incident Management โ€ข Develop actionable alerts that identify meaningful ... SRE, cloud, Kubernetes, observability, and incident-management practices. Required Knowledge ...

DevOps & Site Reliability Engineer

Houston, TX

$47.25 - $62.75/hr

... SRE ENGINEER Location: HOUSTON, TX FLSA Class: EXEMPT Responsible to: Directo of Software ... Implement and maintain monitoring, alerting, and observability systems (Prometheus, Grafana ...

DevOps & Site Reliability Engineer

Houston, TX ยท On-site

$47.25 - $62.75/hr

... SRE ENGINEER Location: HOUSTON, TX FLSA Class: EXEMPT Responsible to: Directo of Software ... Implement and maintain monitoring, alerting, and observability systems (Prometheus, Grafana ...

Site Reliability Engineer III

Houston, TX

$54.50 - $72.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate ... Experience in observability such as white and black box monitoring, service level objective ...

Site Reliability Engineer III

Houston, TX

$54.50 - $72.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk ... Experience in observability such as white and black box monitoring, service level objective ...

Site Reliability Engineer III

Houston, TX ยท On-site

$54.50 - $72.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate ... Experience in observability such as white and black box monitoring, service level objective ...

Site Reliability Engineer III

Houston, TX

$54.50 - $72.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate ... Experience in observability such as white and black box monitoring, service level objective ...

next page

Showing results 1-20

Observability Site Reliability Engineer information

See Houston, TX salary details

$10

$60

$87

How much do observability site reliability engineer jobs pay per hour?

As of Aug 27, 2026, the average hourly pay for observability site reliability engineer in Houston, TX is $60.87, according to ZipRecruiter salary data. Most workers in this role earn between $52.36 and $69.57 per hour, depending on experience, location, and employer.

What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?

AspectObservability Site Reliability EngineerMonitoring Engineer
FocusEnsuring system reliability through observability, automation, and incident responseImplementing and managing monitoring tools and dashboards
SkillsCloud platforms, scripting, incident management, observability toolsMonitoring tools, alerting systems, data analysis
Work EnvironmentDevOps teams, cloud infrastructure, large-scale systemsOperations teams, infrastructure monitoring

While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.

What are popular job titles related to Observability Site Reliability Engineer jobs in Houston, TX?

For Observability Site Reliability Engineer jobs in Houston, TX, the most frequently searched job titles are:

What job categories do people searching Observability Site Reliability Engineer jobs in Houston, TX look for?

The top searched job categories for Observability Site Reliability Engineer jobs in Houston, TX are:

What cities near Houston, TX are hiring for Observability Site Reliability Engineer jobs?

Cities near Houston, TX with the most Observability Site Reliability Engineer job openings:

Staff Observability Platform Engineer (SRE)

Recruitment.ai

Houston, TX โ€ข On-site

$54.50 - $72.25/hr

Other

Posted 8 days ago


Job description

Staff Observability Platform Engineer (SRE)
Locations: Seattle, WA (Hybrid), Houston, TX (Hybrid), New York, NY (Hybrid)
What Youโ€™ll Do
  • Design, build, and evolve observability platforms across metrics, logs, traces, alerting, and telemetry pipelines.
  • Lead the implementation of scalable observability solutions that support Nscaleโ€™s growing GPU and AI infrastructure.
  • Partner with SRE, infrastructure, platform, and AI/ML teams to ensure observability is embedded throughout the software and infrastructure lifecycle.
  • Drive improvements in monitoring coverage, alert quality, service health visibility, and incident response effectiveness.
  • Develop standards, frameworks, and reusable patterns that simplify observability adoption across engineering teams.
  • Identify reliability risks and operational blind spots, helping teams proactively address them before they impact customers.
  • Contribute to architectural decisions around telemetry collection, storage, retention, cardinality management, and performance optimization.
  • Lead technical initiatives and projects that improve platform scalability, reliability, and operational efficiency.
  • Mentor engineers and provide technical guidance through design reviews, code reviews, and knowledge sharing.
  • Participate in incident investigations and postmortems, translating operational learnings into durable platform improvements.
  • Evaluate new observability technologies and practices, balancing innovation with operational simplicity and long-term maintainability.
About You
  • Experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines.
  • Strong experience building and operating observability platforms in cloud-native, distributed environments.
  • Deep hands-on experience with several of the following technologies: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic, or similar platforms.
  • Strong software engineering skills with proficiency in Go, Python, or equivalent languages.
  • Experience operating and troubleshooting Kubernetes-based platforms at scale.
  • Strong understanding of monitoring, logging, tracing, telemetry pipelines, and modern observability practices.
  • Experience designing systems with scalability, reliability, performance, and operational simplicity in mind.
  • Proficiency with Infrastructure-as-Code tools such as Terraform, Ansible, or equivalent.
  • Ability to lead technical initiatives and influence engineering decisions across multiple teams.
  • Excellent communication skills with the ability to explain technical tradeoffs and align stakeholders around pragmatic solutions.
Preferred
  • Experience operating observability systems in GPU, AI/ML, HPC, or large-scale compute environments.
  • Familiarity with Slurm, Kubernetes GPU scheduling, or AI infrastructure platforms.
  • Experience with high-volume telemetry pipelines and streaming technologies such as Kafka, Vector, or Fluent Bit.
  • Knowledge of observability challenges related to model training, inference workloads, GPU utilization, and distributed AI systems.
  • Experience mentoring engineers and helping grow technical capability across teams.