Site Reliability Engineer NEX
$54.50 - $72.25/hr
The Site Reliability Engineer will partner with software engineering, data engineering, and ... Observability and Incident Management Develop actionable alerts that identify meaningful service ...
$54.50 - $72.25/hr
The Site Reliability Engineer will partner with software engineering, data engineering, and ... Observability and Incident Management Develop actionable alerts that identify meaningful service ...
$54.50 - $72.25/hr
The Site Reliability Engineer will partner with software engineering, data engineering, and ... Observability and Incident Management Develop actionable alerts that identify meaningful service ...
Houston, TX · On-site
$54.50 - $72.25/hr
Monitor systems using observability tools (metrics, logs, tracing) and respond to incidents ... Site Reliability Engineering, DevOps, or similar roles * Strong experience with Google Cloud ...
Houston, TX · On-site
$54.50 - $72.25/hr
Monitor systems using observability tools (metrics, logs, tracing) and respond to incidents ... Site Reliability Engineering, DevOps, or similar roles * Strong experience with Google Cloud ...
Spring, TX · On-site
$155K - $222K/yr
Meet the Team The SRE Fleet team is responsible for maintaining the stability, scalability, and ... Knowledge of monitoring, observability, and reliability engineering practices and tooling.
Spring, TX · On-site
$155K - $222K/yr
Meet the Team The SRE Fleet team is responsible for maintaining the stability, scalability, and ... Knowledge of monitoring, observability, and reliability engineering practices and tooling.
Houston, TX · On-site
$155K - $222K/yr
Meet the Team The SRE Fleet team is responsible for maintaining the stability, scalability, and ... Knowledge of monitoring, observability, and reliability engineering practices and tooling.
Houston, TX · On-site
$155K - $222K/yr
Meet the Team The SRE Fleet team is responsible for maintaining the stability, scalability, and ... Knowledge of monitoring, observability, and reliability engineering practices and tooling.
$54.50 - $72.25/hr
As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk ... Experience in observability such as white and black box monitoring, service level objective ...
$54.50 - $72.25/hr
As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk ... Experience in observability such as white and black box monitoring, service level objective ...
Houston, TX · On-site
$54.50 - $72.25/hr
The Senior Site Reliability Engineer is responsible for improving the reliability, availability ... Strongproficiencywith observability platforms (e.g., Datadog, Prometheus/Grafana, ELK/OpenSearch ...
Houston, TX · On-site
$54.50 - $72.25/hr
The Senior Site Reliability Engineer is responsible for improving the reliability, availability ... Strongproficiencywith observability platforms (e.g., Datadog, Prometheus/Grafana, ELK/OpenSearch ...
Houston, TX · On-site
$54.50 - $72.25/hr
Site Reliability Engineer (SRE) - Microsoft Hyper-V & Private Cloud Location: Jersey City, NJ / Houston, TX / San Francisco, CA Job Type: Contract Experience: Minimum 8+ Years Required About the Role ...
Quick apply
Houston, TX · On-site
$54.50 - $72.25/hr
Site Reliability Engineer (SRE) - Microsoft Hyper-V & Private Cloud Location: Jersey City, NJ / Houston, TX / San Francisco, CA Job Type: Contract Experience: Minimum 8+ Years Required About the Role ...
Houston, TX · On-site
$54.50 - $72.25/hr
Job Title: SRE Engineer Location: Houston, TX and Jersey City, NJ - 3 Days Onsite Role FTE role with Mphasis Client: Mphasis H1B transfer will work for this role. No, JNTU and Osmania candidate. Role ...
Houston, TX · On-site
$54.50 - $72.25/hr
Job Title: SRE Engineer Location: Houston, TX and Jersey City, NJ - 3 Days Onsite Role FTE role with Mphasis Client: Mphasis H1B transfer will work for this role. No, JNTU and Osmania candidate. Role ...
Houston, TX · On-site
$54.50 - $72.25/hr
As a Lead Site Reliability Engineer at JPMorgan Chase within Market Risk Technology, you hold a ... Proficiency and experience in observability such as white and black box monitoring, SLO alerting ...
Houston, TX · On-site
$54.50 - $72.25/hr
As a Lead Site Reliability Engineer at JPMorgan Chase within Market Risk Technology, you hold a ... Proficiency and experience in observability such as white and black box monitoring, SLO alerting ...
Houston, TX · On-site
$54.50 - $72.25/hr
Exempt Summary The SRE Developer is responsible for supporting, maintaining, and improving production critical systems to ensure reliability, availability, and performance within a manufacturing ...
Quick apply
Houston, TX · On-site
$54.50 - $72.25/hr
Exempt Summary The SRE Developer is responsible for supporting, maintaining, and improving production critical systems to ensure reliability, availability, and performance within a manufacturing ...
Houston, TX · On-site
$54.50 - $72.25/hr
Exempt Summary The SRE Developer is responsible for supporting, maintaining, and improving production critical systems to ensure reliability, availability, and performance within a manufacturing ...
Houston, TX · On-site
$54.50 - $72.25/hr
Exempt Summary The SRE Developer is responsible for supporting, maintaining, and improving production critical systems to ensure reliability, availability, and performance within a manufacturing ...
Houston, TX · On-site
$123K/yr
... Site Reliability Engineering (SRE) • Power BI, Power Query, DAX & SQL Server / Azure SQL • ... Automation, Observability & Performance Optimization • Strong Stakeholder Management ...
Houston, TX · On-site
$123K/yr
... Site Reliability Engineering (SRE) • Power BI, Power Query, DAX & SQL Server / Azure SQL • ... Automation, Observability & Performance Optimization • Strong Stakeholder Management ...
Apply modern engineering capabilities where appropriate (i.e. observability, Site Reliability Engineering (SRE), security engineering, containerization, infrastructureascode, AI/ML integration, and ...
Apply modern engineering capabilities where appropriate (i.e. observability, Site Reliability Engineering (SRE), security engineering, containerization, infrastructureascode, AI/ML integration, and ...
Houston, TX · On-site
$54.50 - $72.25/hr
As a Cloud Platform Engineer (Site Reliability) , you will: * Developing new cloud-native platform ... Observability Monitoring tools such as Grafana and SuperSet * Proficiency with Infrastructure-as ...
Houston, TX · On-site
$54.50 - $72.25/hr
As a Cloud Platform Engineer (Site Reliability) , you will: * Developing new cloud-native platform ... Observability Monitoring tools such as Grafana and SuperSet * Proficiency with Infrastructure-as ...
$54.50 - $72.25/hr
As a Cloud Platform Engineer (Site Reliability) , you will: * Developing new cloud-native platform ... Observability Monitoring tools such as Grafana and SuperSet * Proficiency with Infrastructure-as ...
$54.50 - $72.25/hr
As a Cloud Platform Engineer (Site Reliability) , you will: * Developing new cloud-native platform ... Observability Monitoring tools such as Grafana and SuperSet * Proficiency with Infrastructure-as ...
Houston, TX · On-site
$124K - $159K/yr
Implement observability solutions using metrics, logging, and tracing to enable proactive issue ... Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity ...
Houston, TX · On-site
$124K - $159K/yr
Implement observability solutions using metrics, logging, and tracing to enable proactive issue ... Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity ...
Houston, TX · On-site
$124K - $159K/yr
Implement observability solutions using metrics, logging, and tracing to enable proactive issue ... Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity ...
Houston, TX · On-site
$124K - $159K/yr
Implement observability solutions using metrics, logging, and tracing to enable proactive issue ... Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity ...
Houston, TX · On-site
$124K - $159K/yr
Implement observability solutions using metrics, logging, and tracing to enable proactive issue ... Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity ...
Houston, TX · On-site
$124K - $159K/yr
Implement observability solutions using metrics, logging, and tracing to enable proactive issue ... Apply Site Reliability Engineering (SRE) practices such as SLIs/SLOs, error budgets, capacity ...
Houston, TX · On-site
$241K/yr
Infrastructure Development & SRE Team Buildout * Staff, train, and mature a new team that becomes ... Observability platforms (metrics, logs, traces, SLOs, alerts) and incident management practices.
Houston, TX · On-site
$241K/yr
Infrastructure Development & SRE Team Buildout * Staff, train, and mature a new team that becomes ... Observability platforms (metrics, logs, traces, SLOs, alerts) and incident management practices.
Houston, TX · On-site
$241K/yr
Infrastructure Development & SRE Team Buildout * Staff, train, and mature a new team that becomes ... Observability platforms (metrics, logs, traces, SLOs, alerts) and incident management practices.
Houston, TX · On-site
$241K/yr
Infrastructure Development & SRE Team Buildout * Staff, train, and mature a new team that becomes ... Observability platforms (metrics, logs, traces, SLOs, alerts) and incident management practices.
$9.26 - $15.57
1% of jobs
$15.57 - $21.87
0% of jobs
$21.87 - $28.18
0% of jobs
$28.18 - $34.48
2% of jobs
$34.48 - $40.79
4% of jobs
$40.79 - $47.09
17% of jobs
$47.15 is the 25th percentile. Wages below this are outliers.
$47.09 - $53.40
27% of jobs
$53.40 - $59.70
20% of jobs
$60.78 is the 75th percentile. Wages above this are outliers.
$59.70 - $66.01
17% of jobs
$66.01 - $72.31
6% of jobs
$72.31 - $78.62
4% of jobs
$9
$54
$78
| Aspect | Observability Site Reliability Engineer | Monitoring Engineer |
|---|---|---|
| Focus | Ensuring system reliability through observability, automation, and incident response | Implementing and managing monitoring tools and dashboards |
| Skills | Cloud platforms, scripting, incident management, observability tools | Monitoring tools, alerting systems, data analysis |
| Work Environment | DevOps teams, cloud infrastructure, large-scale systems | Operations teams, infrastructure monitoring |
While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.
5.0
Based on 23 frontline employees who took The Breakroom Quiz
80th of 86 rated oil and gas companies
This role combines software engineering, cloud infrastructure, automation, and production operations. The Site Reliability Engineer will partner with software engineering, data engineering, and platform teams to improve service reliability, reduce operational toil, strengthen production readiness, and establish consistent reliability practices across the technology organization. The successful candidate will have experience operating production systems, automating infrastructure and operational processes, improving observability, and applying Site Reliability Engineering principles such as service-level indicators, service-level objectives, error budgets, and blameless postmortems.
Responsibilities Reliability Engineering
Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform.
Improve service availability, latency, performance, scalability, and operational resilience.
Define, implement, and track service-level indicators, service-level objectives, and error budgets.
Perform capacity planning, performance analysis, and workload forecasting.
Design and validate disaster recovery, backup, failover, and service-restoration capabilities.
Implement and maintain secure cloud networking, IAM, workload identities, service accounts, and access-control practices.
Partner with cybersecurity and identity teams to ensure infrastructure and services follow organizational security standards.
Monitor cloud consumption and optimize resource utilization, performance, and cost efficiency.
Identify operational risks and recommend improvements to cloud architecture and service design.
Automation and Platform Engineering
Build and maintain cloud infrastructure using Terraform or comparable infrastructure-as-code tools.
Automate repetitive operational activities and systematically identify, measure, and reduce manual toil.
Build reusable infrastructure modules, deployment patterns, and operational tooling.
Improve CI/CD pipelines to enable secure, repeatable, and reliable software delivery.
Observability and Incident Management
Develop actionable alerts that identify meaningful service degradation while reducing alert fatigue and unnecessary operational noise.
Create and maintain dashboards, runbooks, operational procedures, and troubleshooting documentation.
Participate in a sustainable on-call rotation supporting production systems.
Respond to production incidents, coordinate service restoration, and lead incident response when appropriate.
Facilitate blameless postmortems and identify corrective and preventive actions.
Use incident and operational data to improve system design, automation, monitoring, and response processes.
Collaboration and Service Ownership
Partner with software engineering, data engineering, security, and product teams to improve application reliability and production operations.
Promote shared responsibility for production reliability between application development and platform teams.
Establish and document reliability standards, operational practices, and reusable engineering patterns.
Provide technical guidance and coaching on SRE, cloud, Kubernetes, observability, and incident-management practices.
Required Knowledge, Skills, and Abilities
Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud engineering, production software engineering, or a similar role. Experience operating highly available systems in a 24/7 production environment.
Hands-on experience operating workloads on Google Cloud Platform or another major public cloud platform.
Strong experience managing compute, networking and data GCP services workloads
Strong experience with containerization and orchestration technologies, including Docker and Kubernetes.
Experience building and managing infrastructure with Terraform or a comparable infrastructure-as-code tool.
Proficiency in Python, Go, Java, or another comparable programming language.
Experience implementing or operating CI/CD pipelines using GitHub Actions,Azure DevOps, Bitbucket Pipelines, or comparable tools.
Experience implementing observability using metrics, logs, traces, dashboards, and alerts.
Experience participating in on-call rotations, responding to incidents, and contributing to postmortems.
Understanding of SLIs, SLOs, error budgets, and other SRE principles.
Ability to troubleshoot complex issues across application, infrastructure, network, data, and cloud-service layers.
Ability to communicate effectively with engineering teams, business stakeholders, and operational personnel.
Minimum Qualifications
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
3+ years of experience in Site Reliability Engineering, platform engineering, cloud engineering, or DevOps.
3+ years of experience operating production workloads in GCP.
Ability to understand and communicate in English at a level sufficient to issue, receive, and respond to safety-related and operations-related instructions.
Preferred Qualifications
Google Cloud and/or Kubernetes certifications.
Experience supporting data-intensive, streaming, analytics, or event-driven platforms.
Experience establishing production-readiness, incident-management, or reliability-review processes.
Experience working in the energy, oil and gas, industrial, IoT, field operations, or other operationally critical industries.
Experience supporting technology environments that integrate cloud platforms with remote sites, field equipment, industrial systems, or edge computing.
Patterson-UTI is committed to a workplace free from discrimination and harassment, offering equal employment opportunities to all individuals regardless of personal characteristics protected by law. Employees are encouraged to report any concerns through multiple channels.
Get the full story on Breakroom
Sourced by ZipRecruiter
Oil and gas extraction
5,001 - 10,000 Employees
Houston, TX, US
1978