2

Site Reliability Engineer Remote Jobs in Austin, TX

K8 Site Reliability SME

Austin, TX ยท On-site +1

$180K - $260K/yr

Strong SRE background: SLI/SLO frameworks, incident management, capacity planning * Experience with Prometheus, Grafana, and alerting at scale * Strong programming skills in Go or Python for operator ...

Data Engineer, Hardware Reliability (Starlink)

Bastrop, TX ยท On-site +1

$56.75 - $75.25/hr

DATA ENGINEER, HARDWARE RELIABILITY (STARLINK) SpaceX is leveraging our experience building rockets ... This position is based in Bastrop, TX (Austin area) and requires being on-site; hybrid or remote ...

DevOps Engineer - Remote

Austin, TX ยท Remote

$50 - $150/hr

Remote Job Overview We are seeking experienced DevOps Engineers to create and evaluate production ... Infrastructure as Code Required Qualifications * 3+ years of professional experience in DevOps, SRE ...

AI Cloud Senior DevOps Engineer

Austin, TX ยท On-site +1

$145K - $260K/yr

Bachelor's degree or above in Computer Science, Engineering, or a related technical field, with 5+ years of hands-on experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure ...

Head of Commissioning and Reliability

Austin, TX ยท Remote

$58.25 - $77.50/hr

... remote culture. About This Role As the Head of Commissioning and Reliability on Intersect ... You'll sit at the critical nexus of Engineering, Construction, and Operations, building the systems ...

Showing results 21-40

Site Reliability Engineer Remote information

See Austin, TX salary details

$10

$63

$91

How much do site reliability engineer remote jobs pay per hour?

As of Sep 4, 2026, the average hourly pay for site reliability engineer remote in Austin, TX is $63.18, according to ZipRecruiter salary data. Most workers in this role earn between $54.33 and $72.21 per hour, depending on experience, location, and employer.

What is a site reliability engineer remote?

A Site Reliability Engineer (SRE) in a remote role is responsible for ensuring the reliability, performance, and scalability of software systems while working from a remote location. They bridge the gap between development and operations by implementing automation, monitoring, and incident response strategies. Remote SREs collaborate with distributed teams to improve infrastructure, troubleshoot issues, and optimize system performance. Strong communication skills, proficiency in cloud technologies, and expertise in software development are essential for success in this role.

What are the key skills and qualifications needed to thrive as a site reliability engineer remote?

To thrive as a Site Reliability Engineer Remote, you need expertise in systems administration, cloud infrastructure, automation, coding (often in Python or Go), and a solid grasp of networking fundamentals, usually demonstrated with a degree in computer science or equivalent experience. Familiarity with tools such as Docker, Kubernetes, AWS/GCP/Azure, monitoring platforms like Prometheus, and certifications like AWS Certified SysOps Administrator are highly valued. Excellent problem-solving, communication, and collaboration skills are essential, especially when troubleshooting incidents and passing information across distributed teams. These abilities ensure reliable, scalable services and smooth coordination in a remote work environment.

What are some common challenges faced by site reliability engineers working remotely, and how are they addressed?

Site Reliability Engineers working remotely may encounter challenges like coordinating across multiple time zones, maintaining clear communication during urgent incidents, and managing complex systems without direct on-site access. These are often addressed by leveraging collaborative tools (like Slack, Zoom, and incident management platforms), implementing well-documented processes, and participating in regular team syncs or on-call rotations. Remote SREs also benefit from automation and observability practices that provide in-depth systems insights without needing physical presence. Many organizations support their success through robust onboarding, continuous training, and establishing clear lines of communication for rapid response scenarios. This blend of technical and teamwork strategies helps remote SREs maintain service reliability and stay connected with their colleagues.

What are the most commonly searched types of Site Reliability Engineer jobs in Austin, TX?

The most popular types of Site Reliability Engineer jobs in Austin, TX are:

What are popular job titles related to Site Reliability Engineer Remote jobs in Austin, TX?

For Site Reliability Engineer Remote jobs in Austin, TX, the most frequently searched job titles are:

What job categories do people searching Site Reliability Engineer Remote jobs in Austin, TX look for?

The top searched job categories for Site Reliability Engineer Remote jobs in Austin, TX are:

What cities near Austin, TX are hiring for Site Reliability Engineer Remote jobs?

Cities near Austin, TX with the most Site Reliability Engineer Remote job openings:

Infographic showing various Site Reliability Engineer Remote job openings in Austin, TX as of August 2026, with employment types broken down into 1% As Needed, 83% Full Time, 12% Part Time, 1% Temporary, 2% Contract, and 1% Nights. Highlights an 94% Physical, 2% Hybrid, and 4% Remote job distribution, with an average salary of $131,418 per year, or $63.2 per hour.

Sr SRE & Automation Engineer (Customer Facing)

BitDeer

Austin, TX โ€ข On-site, Remote

$180K - $260K/yr

Full-time

Posted 7 days ago


Job description

About Bitdeer Technologies Group
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/
Job Description
NeoCloud is building an AI-operated GPU cloud - and because it is a customer-facing cloud service, reliability is the product. Tenants run mission-critical training, fine-tuning, and inference workloads on our GPU infrastructure and trust us with their SLAs. In this role you own the reliability of the customer-facing GPU cloud service end-to-end: from tenant onboarding and service provisioning, through workload execution, incident response, and post-incident recovery. You are the SRE who stands between raw infrastructure and the customer's experience - designing the observability, automation, and operational practices that make a 10,000-GPU cloud feel simple and dependable to the tenants who depend on it.
What You'll Own
  • End-to-end reliability of the customer-facing GPU cloud service - availability, job completion, provisioning latency, and tenant experience.
  • Production Kubernetes clusters optimized for GPU workloads at scale (100-10,000 GPUs) as the runtime substrate for customer workloads.
  • Nvidia GPU operator, device plugin, MIG configuration, GPU time-slicing, and multi-tenant GPU allocation policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity - placing customer jobs on the right hardware.
  • Customer & tenant lifecycle: onboarding, quota management, isolation enforcement (namespaces, network policies, RBAC, resource quotas, pod security), and offboarding/reclamation.
  • Bare-Metal-as-a-Service (BMaaS): automated provisioning, tenant handoff, lifecycle, and reclamation.
  • SLIs/SLOs/SLAs for the customer cloud service: cluster availability, job completion rates, provisioning latency, API availability.
  • Incident management with customer communication: runbook automation, escalation, customer-facing status updates, and post-incident reviews.
  • Monitoring & observability stack: Prometheus, Grafana, Alertmanager, PagerDuty - tenant-aware dashboards and alerting.
  • GPU node failure handling: automated detection, drain/cordon/taint, and workload rescheduling - minimizing customer-visible impact.
  • Infrastructure-as-code: Terraform providers/modules, Helm, and GitOps (ArgoCD/Flux) across GPU clusters.
  • Customer-facing operational readiness: service documentation, tenant runbooks, capacity planning, and support tiering.

Customer-Facing Ownership
  • You are accountable for the customer's reliability experience - when a tenant's job fails or a node drops, you own the detection, remediation, and communication loop.
  • Define and publish customer-facing SLAs/SLOs and drive error-budget-based prioritization between feature work and reliability.
  • Partner with customer success / support to close the feedback loop between customer-reported issues and systemic improvements.
  • Build self-service observability that lets customers answer their own questions - status, quota, job health - reducing support load.

Feed the AIOps Substrate
  • The remediation-actuator and workflow engine land here - you make the control plane safe for automated action.
  • Your CRDs and runbooks are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter - turning customer-impacting incidents into self-healing events.

What Success Looks Like in Year 1
  • Customer-facing GPU cloud service SLAs published and met - availability, job completion, provisioning latency.
  • Automated drain/reschedule around predicted GPU faults, at scale, without customer-visible impact.
  • BMaaS live for external tenants with self-service onboarding.
  • MTTD and MTTR for customer-impacting incidents reduced through automation.
  • Tenant self-service observability live - customers can see their own job health, quota, and status.

Requirements
  • 5+ years in SRE / cloud operations, with at least 2 years operating GPU workloads at scale.
  • Deep understanding of Kubernetes operations and GPU workload management (Nvidia GPU operator, device plugin, MIG, time-slicing, GPU scheduling).
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
  • Customer-facing cloud service experience - defining and operating against customer SLAs/SLOs, handling tenant incidents and communications.
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
  • Strong SRE background: SLI/SLO/SLA frameworks, error budgets, incident management, capacity planning.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for automation / operator development.
  • AIOps aptitude - you view the control plane as an execution surface for automated remediation, not just a scheduler.
  • Runbook-as-code mindset - every SRE playbook you write should be executable by the platform.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.