2

Remote Reliability Engineer Jobs in Grand Prairie, TX

DevOps Expert - Remote

Dallas, TX ยท Remote

$30 - $80/hr

Remote Job Overview We are seeking experienced Developer & Infrastructure Experts to evaluate AI ... SRE, and platform engineering. You will test AI-generated commands, configurations, and workflows ...

Dynatrace Engineer

Dallas, TX ยท Remote

$58.25 - $77.50/hr

Remote The Mission ConglomerateIT needs a Dynatrace Engineer to take full ownership of ... Familiarity with Site Reliability Engineering (SRE) principles The Profile We're After A hands-on ...

Remote Job Overview We are seeking experienced Developer & Infrastructure Experts to evaluate AI-powered workflows across software development, cloud infrastructure, DevOps, SRE, and platform ...

Sr. Software Engineer (Java)

Irving, TX ยท Remote

$55 - $58/hr

... , and Application teams to modernize and secure homegrown applications while supporting automation and cloud initiatives. Work Arrangement * Remote: Candidates must be able to work on EST hours.

Sr. DevOps Engineer

Dallas, TX ยท Remote

$128K - $165K/yr

You will be responsible for owning and evolving our CI/CD infrastructure, driving SRE best ... Proven ability to work effectively in remote teams across time zones. Preferred Qualifications

Showing results 21-40

Remote Reliability Engineer information

See Grand Prairie, TX salary details

$57.7K

$111.7K

$133.5K

How much do remote reliability engineer jobs pay per year?

As of Sep 5, 2026, the average yearly pay for remote reliability engineer in Grand Prairie, TX is $111,668.00, according to ZipRecruiter salary data. Most workers in this role earn between $97,000.00 and $122,100.00 per year, depending on experience, location, and employer.

What is a remote reliability engineer?

A Remote Reliability Engineer is a professional who works from a remote location to ensure that systems, applications, or infrastructure are reliable, available, and performing well. Their responsibilities typically include monitoring system health, diagnosing issues, implementing preventative measures, and collaborating with teams to improve system reliability. They often use tools for automation, incident response, and performance monitoring, all while working offsite. This role is critical in minimizing downtime and ensuring a smooth user experience, especially for companies with complex technical environments. Remote Reliability Engineers must have strong problem-solving skills and be proficient in cloud technologies, automation, and incident management.

What are the key skills and qualifications needed to thrive as a remote reliability engineer?

To thrive as a Remote Reliability Engineer, you need a strong background in systems engineering, software development, and infrastructure management, often supported by a degree in computer science or a related field. Proficiency with cloud platforms (such as AWS, Azure, or GCP), monitoring tools (like Prometheus, Grafana), and relevant certifications (e.g., AWS Certified DevOps Engineer) is highly valuable. Excellent problem-solving, communication, and collaboration skills are crucial for working effectively across distributed teams and responding to incidents. These abilities ensure system reliability, quick incident resolution, and seamless remote teamwork, which are vital for maintaining high service uptime and user satisfaction.

How do remote reliability engineers typically collaborate with on-site teams to address urgent technical issues?

Remote Reliability Engineers often utilize a combination of video conferencing, instant messaging, and collaborative monitoring tools to stay closely connected with on-site teams. When urgent technical issues arise, they participate in real-time troubleshooting sessions, analyze system logs remotely, and may guide on-site staff through step-by-step resolution procedures. Building strong communication channels and regular check-ins are essential to ensure swift and effective collaboration, even across different time zones. This structure allows Remote Reliability Engineers to contribute significantly to system uptime while working from a distance.

What is the difference between Remote Reliability Engineer vs Remote Site Reliability Engineer?

AspectRemote Reliability EngineerRemote Site Reliability Engineer
CredentialsTypically requires certifications like AWS Certified Solutions Architect, Linux Foundation certificationsSimilar credentials, often with additional focus on site-specific tools and monitoring
Work EnvironmentPrimarily remote, focusing on cloud infrastructure and system reliabilityRemote with some on-site responsibilities, focusing on infrastructure and operational stability
Industry UsageUsed across tech, cloud providers, SaaS companiesCommon in data centers, cloud providers, and large enterprise IT
Search & Comparison IntentOften compared due to overlapping roles in system reliability and cloud infrastructureCompared for on-site vs remote operational responsibilities

The main difference is that Remote Reliability Engineers focus on cloud and system reliability remotely, while Remote Site Reliability Engineers may have some on-site duties related to infrastructure. Both roles require similar skills and certifications but differ in their work environment and specific responsibilities.

What are the most commonly searched types of Reliability Engineer jobs in Grand Prairie, TX?

The most popular types of Reliability Engineer jobs in Grand Prairie, TX are:

What are popular job titles related to Remote Reliability Engineer jobs in Grand Prairie, TX?

For Remote Reliability Engineer jobs in Grand Prairie, TX, the most frequently searched job titles are:

What cities near Grand Prairie, TX are hiring for Remote Reliability Engineer jobs?

Cities near Grand Prairie, TX with the most Remote Reliability Engineer job openings:

Infographic showing various Remote Reliability Engineer job openings in Grand Prairie, TX as of August 2026, with employment types broken down into 92% Full Time, 2% Part Time, and 6% Contract. Highlights an 87% Physical, 4% Hybrid, and 9% Remote job distribution, with an average salary of $111,668 per year, or $53.7 per hour.

Sr. Site Reliability Engineer (Storage Platform)

3B Staffing LLC

Dallas, TX โ€ข Remote

$66/hr

Full-time, Contractor

This job post hasย expired today.ย Applications are no longer accepted.


Job description

Job Title: Sr. Site Reliability Engineer (Storage Platform)

Client: AHEAD

VISA: US Citizens / GC

Salary: $66/Hr. on C2C
Job Type: Contract, with potential for Contract-to-hire - the client only wants to see candidates that are willing to convert to full-time employment for this role that do not require any type of sponsorship.
Worksite Requirement: Fully Remote
Industry: IT Services and IT Consulting
Interview Process: 2-3 rounds of video conference interviews

Job Summary

We are seeking a highly experienced Sr Site Reliability Engineer - Storage Platforms to design, implement, and support Software Defined Storage (SDS) and Kubernetes platforms in a private cloud environment. This role focuses on scalability, resilience, automation, and performance using Infrastructure-as-Code and GitOps practices.

This is a deeply technical role requiring expert-level understanding of Software Defined Storage, Kubernetes, and extensive working knowledge on Linux Operating systems. You will also collaborate with platform and SRE teams to maintain secure, performant, and multitenant-isolated services that serve high-throughput, mission-critical applications.

Key Responsibilities

  • Design, implement, and operate large-scale Software Defined Storage architectures across private and public cloud regions within ITIL methodology.
  • Deploy and support enterprise storage platforms (Pure Storage, HPE, NetApp) and SDS solutions (Ceph, Longhorn).
  • Build self-service storage workflows for Kubernetes CSI and OpenStack consumers (VM and Baremetal).
  • Develop Infrastructure-as-Code using Ansible, Terraform, Helm and Git, with Python/Bash automation.
  • Implement CI/CD pipelines for infrastructure updates, patching, upgrades, testing, and rollback.
  • Build observability, alerting, and auto-remediation using GitOps and tools such as Prometheus, Loki, and Grafana.
  • Architect and maintain high availability, disaster recovery, and scale-out infrastructure.
  • Develop and review high-level and low-level design documents for storage infrastructure
  • Perform deep troubleshooting across storage, Kubernetes, hypervisors, networking, and Linux systems.
  • Participate in on-call rotations, incident response, and root cause analysis.
  • Collaborate globally on change management, documentation, and operational best practices.

Must Have

  • 6+ years of experience managing enterprise storage and Kubernetes platforms on Linux.
  • Strong hands-on experience with SDS solutions (Ceph, Longhorn) and storage migrations from legacy systems.
  • Experience with block, file, and object storage, including Fibre Channel and IP-based protocols.
  • Experience with NVMe-oF or iSCSI fabrics.
  • Expert knowledge of Kubernetes and Linux systems (Ubuntu, RHEL/CentOS).
  • Proficiency with Infrastructure-as-Code (IaC) (Ansible, Terraform).
  • Strong scripting skills in Python and Bash (Golang (GO) a plus).
  • Strong working knowledge of Enterprise DNS and integrations with Kubernetes
  • Experience operating 24x7 mission-critical production environments.
  • Hands-on experience with KVM hypervisors (Suse Harvester, OpenStack).
  • Strong written and verbal communication skills.
  • Proficiency with Git, CI/CD pipelines, and automated testing frameworks
  • Ability to write technical documentation and contribute to community wikis or knowledge bases.
  • Bachelor's degree in computer science or equivalent professional experience.

Nice to Have

  • OpenStack Cinder multi-backend administration.
  • Backup platforms (Rubrik).
  • Understanding of CIS/NIST security and infrastructure lifecycle management.
  • ITIL Foundation/advanced certifications in support of ITSM standard methodology.
  • Background in telco, edge cloud, or large enterprise environments.
  • CNCF Certified Kubernetes Administrator (CKA), Certified Kubernetes Security
  • Specialist (CKS) or Red Hat specialist in Ceph Storage Administrator (EX125) certifications.
  • Master's degree in computer science, IT, Engineering, or a related field preferred;
  • equivalent experience and relevant industry certifications will also be considered.