1

Site Reliability Engineering Jobs (NOW HIRING)

Site Reliability Engineering

Los Angeles, CA ยท On-site

$61.50 - $81.50/hr

Site Reliability Engineering (SRE) Location: Los Angeles, CA Remote position Fulltime position JD * Site Reliability Engineer * Experience in Cloud platforms (AWS, Azure, Google Cloud) and hybrid ...

Director, Site Reliability Engineering

Frisco, TX ยท On-site

$158.90 - $295.10/hr

As Director of Site Reliability Engineering, you will lead a global organization responsible for the reliability, scalability, and operational excellence of critical applications and services.

New

The SRE team operates as a guiding and consultative partner to application and development teams rather than owning systems directly. Success in this role requires technical credibility, strong ...

New

Infosec Site Reliability Engineering

Timonium, MD ยท On-site

$54.75 - $72.75/hr

Must Haves: 4+ years experience SRE 2+ years experience with monitoring and observability platforms (NewRelic, Prometheus, Grafana, etc.) 2+ years Cloud Experience, Chaos Engineering experience ...

New

SDF is looking for a Director of Site Reliability Engineering to lead a small, high-leverage SRE team and help shape how engineering teams own, operate, and improve production services. This is a ...

New

Director, Site Reliability Engineering

$58.25 - $77.50/hr

We're looking for a Senior Manager of Site Reliability Engineering to join our team. You'll lead a team of ~10 SREs across North America, UK, HK, and New Zealand - owning both the day-to-day ...

next page

Showing results 1-20

Site Reliability Engineering information

See salary details

$10

$63

$91

How much do site reliability engineering jobs pay per hour?

As of Aug 8, 2026, the average hourly pay for site reliability engineering in the United States is $63.74, according to ZipRecruiter salary data. Most workers in this role earn between $54.81 and $72.84 per hour, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as a site reliability engineer, and why are they important?

To thrive as a Site Reliability Engineer, you need a solid background in software engineering, systems administration, and troubleshooting, often supported by a degree in computer science or related field. Familiarity with automation tools, cloud platforms (such as AWS, GCP, or Azure), containerization (Docker, Kubernetes), and monitoring systems is typically required. Strong problem-solving skills, effective communication, and a proactive mindset set outstanding SREs apart. These skills ensure high system reliability, efficient incident response, and seamless collaboration across development and operations teams.

How does a site reliability engineer typically collaborate with development and operations teams?

Site Reliability Engineers (SREs) work closely with both development and operations teams to ensure systems are reliable, scalable, and efficient. They often participate in code reviews, help define service level objectives (SLOs), and develop automation tools to streamline deployment and incident response. SREs act as a bridge between development and IT, translating operational needs into engineering solutions and vice versa. Regular communication and joint problem-solving are essential parts of the role, fostering a culture of shared responsibility for system uptime and performance.

What is site reliability engineering?

Site Reliability Engineering (SRE) is a discipline that combines aspects of software engineering and IT operations to build and run scalable, reliable, and efficient systems. SRE teams are responsible for ensuring the availability, performance, and reliability of critical services by automating manual processes, monitoring systems, and responding to incidents. They often work closely with development teams to improve system architecture, deploy new features safely, and maintain service-level objectives (SLOs). Overall, SRE aims to create a bridge between development and operations to deliver robust and reliable software services.
More about Site Reliability Engineering jobs
What cities are hiring for Site Reliability Engineering jobs? Cities with the most Site Reliability Engineering job openings:
What are the most commonly searched types of Site Reliability Engineering jobs? The most popular types of Site Reliability Engineering jobs are:
What states have the most Site Reliability Engineering jobs? States with the most job openings for Site Reliability Engineering jobs include:
Infographic showing various Site Reliability Engineering job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 82% Full Time, 12% Part Time, 1% Temporary, and 4% Contract. Highlights an 95% Physical, 2% Hybrid, and 3% Remote job distribution, with an average salary of $132,583 per year, or $63.7 per hour.

Site Reliability Engineering

Sarian, Inc.

Los Angeles, CA โ€ข On-site

$61.50 - $81.50/hr

Full-time

Re-posted 6 days ago


Job description

Role: Site Reliability Engineering (SRE)
Location: Los Angeles, CA
Remote position
Fulltime position
JD
  • Site Reliability Engineer
  • Experience in Cloud platforms (AWS, Azure, Google Cloud) and hybrid environments.
  • Proficiency in container technologies (Docker, Container, Podman).
  • Strong knowledge of Linux administration and networking concepts.
  • Experience with Infrastructure as Code (IaC) tools like Terraform, Ansible, Helm, or Pulumi.
  • Monitoring and logging expertise using Prometheus, Grafana, ELK, Datadog, or Splunk.
  • Hands-on experience with CI/CD pipelines and DevOps tools (Jenkins, GitHub Actions, GitLab CI, ArgoCD).
  • Proficiency in scripting/programming (Python, Bash, Go) for automation.
  • Strong troubleshooting and incident management skills.
  • We are seeking a highly skilled - Site Reliability Engineer (SRE) to manage, optimize, and ensure the reliability of infrastructure.
  • The ideal candidate will have deep expertise in ELK, Dynatrace Pagerduty.
  • Powershell, container orchestration, cloud infrastructure, and automation, along with a strong focus on reliability, scalability, and performance. Good to have Logic Monitor and Python knowledge
  • Reliability & Performance: Implement best practices to ensure high availability, scalability, and performance of containerized applications.
  • Monitoring & Incident Response: Set up monitoring (Prometheus, Grafana, ELK, Dynatrace, Pagerduty, Powershell etc.), troubleshoot issues, and lead incident resolution.
  • Automation & Infrastructure as Code (IaC): Develop and maintain Terraform, Helm charts, and Kubernetes manifests for automation.
  • CI/CD & DevOps Integration: Work with DevOps teams to optimize CI/CD pipelines for Kubernetes deployments (Jenkins, ArgoCD, FluxCD, etc.).
  • Security & Compliance: Implement security best practices for containerized workloads, RBAC, network policies, and vulnerability scanning.
  • Capacity Planning & Optimization: Analyze resource usage and optimize infrastructure costs and performance.
  • Disaster Recovery & Backup: Implement backup and disaster recovery strategies for Kubernetes workloads.

Thanks, and have a nice day
Manikanth
Sarian Solutions, Inc. |Ph: 732-790-2266 x 201 |Fax: 732-696-4242|manikanth.d@sariansolutions.com
www.sariansolutions.com | Certified Minority Business Enterprise (WMBE)
follow us: @sariansol | Check our current openings at https://www.sarianinc.com/work-with-us