2

Remote Hardware Reliability Engineer Jobs in California

Site Reliability Engineer

San Francisco, CA · Remote

$67.25 - $89.25/hr

Remote (US) Department: Cloud Platform Engineering / SRE/Reliability Position summary The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One ...

Senior Site Reliability Engineer

San Diego, CA · Remote

$60.50 - $80.50/hr

This is a remote, contract opportunity for a project Arctiq is delivering for a client. Candidates ... The Senior Site Reliability Engineer is a technical leader responsible for architecting the ...

Cloud DevOps & SRE Engineer

Saratoga, CA · Remote

$62.75 - $83.50/hr

While we embrace a remote-first culture, time zones will be an important consideration for ... What You'll Do * Champion reliability and uptime: Define, measure, and maintain Service Level ...

Cloud DevOps & SRE Engineer

Saratoga, CA · Remote

$62.75 - $83.50/hr

While we embrace a remote-first culture, time zones will be an important consideration for ... Champion reliability and uptime: Define, measure, and maintain Service Level Objectives (SLOs) and ...

$98K - $138K/yr

The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining ... Wellness initiatives #BI-Remote DYN365, Inc d/b/a Restaurant365 is an equal opportunity employer.

$98K - $138K/yr

The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining ... Wellness initiatives #BI-Remote DYN365, Inc d/b/a Restaurant365 is an equal opportunity employer.

Sr. Site Reliability Engineer

San Francisco, CA · On-site +1

$67.25 - $89.25/hr

Open to remote or San Francisco Bay Area, Nashville Metro Area, or Raleigh, NC Area What you\'ll do ... SRE, platform, or staff infrastructure role * Deep Kubernetes expertise across managed (EKS, GKE ...

AI-First SRE/DevOps Engineer

San Jose, CA · On-site +1

$66.75 - $88.75/hr

US (Remote or HQ Hybrid) Job Type: Full-time Axiad is seeking a skilled AI-First SRE/DevOps Engineer with 5-8 years of hands-on infrastructure and platform engineering experience to help build and ...

next page

Showing results 1-20

Remote Hardware Reliability Engineer information

What is a Remote Hardware Reliability Engineer?

A Remote Hardware Reliability Engineer is a professional who assesses, tests, and improves the reliability and durability of hardware components and systems, all while working remotely. They analyze failure data, design tests, and recommend modifications to ensure products meet quality and performance standards. This role often involves collaborating with engineering teams using digital communication tools, running simulations, and reviewing data to prevent future hardware failures. By working remotely, these engineers can support projects from anywhere, offering flexibility while maintaining high standards for hardware reliability.

What is the difference between Remote Hardware Reliability Engineer vs Remote Hardware Test Engineer?

AspectRemote Hardware Reliability EngineerRemote Hardware Test Engineer
CredentialsEngineering degree, certifications in reliability or hardware engineeringEngineering degree, certifications in testing or quality assurance
Work EnvironmentDesigning reliability strategies, analyzing failure data, improving hardware durabilityDeveloping and executing testing procedures, validating hardware performance
Industry UsageManufacturers, tech companies, aerospace, automotiveManufacturers, consumer electronics, hardware development firms

While both roles focus on hardware, the Remote Hardware Reliability Engineer emphasizes ensuring long-term durability and reliability through analysis and design improvements. In contrast, the Remote Hardware Test Engineer concentrates on testing hardware components to verify performance and quality before deployment.

What are the key skills and qualifications needed to thrive as a Remote Hardware Reliability Engineer, and why are they important?

To thrive as a Remote Hardware Reliability Engineer, you typically need a degree in electrical or mechanical engineering, strong analytical skills, and experience with hardware testing and failure analysis. Familiarity with reliability testing tools like HALT/HASS, statistical analysis software, and reliability modeling systems is important, along with certifications such as CRE (Certified Reliability Engineer). Outstanding problem-solving abilities, attention to detail, and effective remote communication are crucial soft skills for collaborating with global teams and addressing issues proactively. These competencies ensure the development and maintenance of robust, reliable hardware products, minimizing failures and optimizing performance in diverse environments.

How does a Remote Hardware Reliability Engineer collaborate effectively with cross-functional teams while working offsite?

As a Remote Hardware Reliability Engineer, effective collaboration with cross-functional teams—such as design, manufacturing, and quality assurance—is typically achieved through regular virtual meetings, collaborative platforms, and detailed documentation. You’ll often participate in remote design reviews, data analysis sessions, and troubleshooting calls to address reliability concerns. Clear communication and proactive project updates are essential to ensure everyone is aligned, especially when working across different time zones or locations. Building strong relationships with team members via consistent online interaction helps maintain project momentum and ensures reliability standards are met.
What are the most commonly searched types of Hardware Reliability Engineer jobs in California? The most popular types of Hardware Reliability Engineer jobs in California are:
What are popular job titles related to Remote Hardware Reliability Engineer jobs in California? For Remote Hardware Reliability Engineer jobs in California, the most frequently searched job titles are:
What job categories do people searching Remote Hardware Reliability Engineer jobs in California look for? The top searched job categories for Remote Hardware Reliability Engineer jobs in California are:
What cities in California are hiring for Remote Hardware Reliability Engineer jobs? Cities in California with the most Remote Hardware Reliability Engineer job openings:
Infographic showing various Remote Hardware Reliability Engineer job openings in California as of July 2026, with employment types broken down into 90% Full Time, 6% Part Time, and 4% Contract. Highlights an 87% Physical, 5% Hybrid, and 8% Remote job distribution.
Site Reliability Engineer

Site Reliability Engineer

STN Inc

San Francisco, CA • Remote

$67.25 - $89.25/hr

Full-time

This job post has expired 1 day ago. Applications are no longer accepted.


Job description

Site Reliability Engineer

Platform and software · shared across customers

Reports to: Director, Site Reliability

Location: Remote (US)

Department: Cloud Platform Engineering / SRE/Reliability

Position summary

The Site Reliability Engineer (SRE) owns reliability, observability, and incident response for the GPU One (GPUaaS) platform. The SRE defines and enforces SLOs aligned with contractual SLAs, builds the observability stack, and leads major incidents to resolution.

Key responsibilities
  • Define and operate Service Level Objectives (SLOs) aligned with customer SLAs

  • Build and maintain the observability stack including metrics, logs, traces, and alerting

  • Lead incident response and chair post-incident reviews

  • Drive automation to reduce toil and improve mean-time-to-recover (MTTR)

  • Author and maintain operational runbooks alongside the NOC

  • Manage on-call rotation, escalation paths, and incident-management tooling

  • Coordinate cross-functionally with NOC, Platform Engineering, and Network Engineering

  • Drive chaos engineering, game days, and reliability testing programs

  • Produce SLA performance reports in coordination with the SLA Manager

  • Mentor junior engineers and contribute to engineering culture

Required qualifications
  • 5+ years in SRE, DevOps, or production engineering roles

  • Strong programming skills in Go, Python, or both

  • Hands-on experience operating Kubernetes-based platforms at scale

  • Deep familiarity with observability tooling (Prometheus, Grafana, Datadog, OpenTelemetry)

  • Strong incident management experience including major-incident command

Preferred qualifications
  • GPU or HPC platform operational experience

  • Familiarity with SLA-driven customer environments and credit calculations

  • Experience with chaos engineering tools (Gremlin, Litmus, or similar)

  • Published SRE content or contributions