1

Product Reliability Engineer Jobs (NOW HIRING)

This team focuses specifically on product reliability engineering for platform hardware and deployed systems - ensuring every Etched product ships with the reliability profile that enterprise and ...

Staff Reliability Engineer

San Diego, CA · On-site

$108K - $137K/yr

This role partners with cross-functional teams spanning technology development, design, product engineering, manufacturing, test, quality, and suppliers to ensure robust product reliability ...

This role partners with cross-functional teams spanning technology development, design, product engineering, manufacturing, test, quality, and suppliers to ensure robust product reliability ...

This role partners with cross-functional teams spanning technology development, design, product engineering, manufacturing, test, quality, and suppliers to ensure robust product reliability ...

... Reliability Engineer to join our innovative team. This role will be in the surgical operating unit ... This role also supports medical device product development, and a solid understanding of design ...

Showing results 41-60

Product Reliability Engineer information

See salary details

$61K

$118K

$141K

How much do product reliability engineer jobs pay per year?

As of Sep 5, 2026, the average yearly pay for product reliability engineer in the United States is $117,973.00, according to ZipRecruiter salary data. Most workers in this role earn between $102,500.00 and $129,000.00 per year, depending on experience, location, and employer.

What is a product reliability engineer?

A Product Reliability Engineer is a professional responsible for ensuring that products consistently perform as intended over their expected lifespan. They analyze product designs, conduct testing, and identify potential failure modes to improve reliability and durability. Their work often involves collaborating with design, manufacturing, and quality teams to implement solutions that enhance product performance and minimize defects. Product Reliability Engineers play a critical role in reducing warranty costs, increasing customer satisfaction, and ensuring compliance with industry standards.

What are the key skills and qualifications needed to thrive as a product reliability engineer?

To thrive as a Product Reliability Engineer, you need a strong background in engineering principles, statistical analysis, and failure mode assessment, often supported by a degree in mechanical, electrical, or related engineering fields. Familiarity with reliability testing tools (such as Weibull++), failure analysis methods, and quality management systems like Six Sigma or FMEA certification is typically required. Exceptional problem-solving, attention to detail, and effective communication skills help you collaborate across teams and convey technical findings clearly. These skills are crucial for ensuring products meet safety, durability, and performance standards, ultimately protecting brand reputation and customer satisfaction.

What are some common challenges product reliability engineers face when working with cross-functional teams?

Product Reliability Engineers often collaborate closely with design, manufacturing, and quality assurance teams to ensure products meet reliability standards. One common challenge is effectively communicating technical reliability requirements to non-engineering stakeholders, which requires translating complex data into actionable insights. Additionally, balancing the need for thorough testing with tight development timelines can create pressure to prioritize tasks strategically. Building strong working relationships and establishing clear processes for feedback and issue resolution can help overcome these challenges.

What is the difference between Product Reliability Engineer vs Software Reliability Engineer?

AspectProduct Reliability EngineerSoftware Reliability Engineer
CredentialsBachelor's in engineering, certifications in reliability or quality assuranceBachelor's in computer science or related field, certifications in software testing or reliability
Work EnvironmentManufacturing, product development, hardware and software integrationSoftware development teams, IT companies, tech firms
Industry UsageManufacturing, electronics, consumer productsTechnology, software services, cloud computing
Job FocusEnsuring product durability, testing hardware/software, improving reliabilityIdentifying software bugs, improving software stability, automating testing

The main difference is that Product Reliability Engineers focus on the overall durability and reliability of physical products and integrated systems, while Software Reliability Engineers concentrate on software stability and bug prevention. Both roles require strong technical skills and collaboration with development teams, but their specific focus areas differ based on the product type.

What does a product reliability engineer do?

A product reliability engineer is responsible for ensuring that products function correctly and consistently over time by analyzing failure data, conducting testing, and implementing improvements. They often use tools like statistical analysis and reliability modeling to identify potential issues and enhance product durability, working closely with design and manufacturing teams. This role typically requires strong problem-solving skills and knowledge of quality standards such as ISO or Six Sigma.
More about Product Reliability Engineer jobs
Infographic showing various Product Reliability Engineer job openings in the United States as of August 2026, with employment types broken down into 79% Full Time, 18% Part Time, 2% Contract, and 1% Nights. Highlights an 89% Physical, 2% Hybrid, and 9% Remote job distribution, with an average salary of $117,973 per year, or $56.7 per hour.

Head of Platform Product Reliability

Delos

San Jose, CA • On-site

$180 - $280/hr

Other

Medical, Dental, Vision

Re-posted 20 hours ago


Key responsibilities

  • Define and own the end‑to‑end reliability strategy for AI servers, accelerator platforms, rack systems and datacenter infrastructure from design requirements through field deployment

  • Lead root‑cause investigations for reliability failures during development, manufacturing and field deployment, driving corrective actions

  • Build fleet reliability infrastructure including telemetry analysis pipelines, field feedback loops and monitoring frameworks


Job description

About Etched

Etched is building hardware for frontier intelligence. We co-design chips, racks, software and manufacturing to deliver best‑in‑class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top‑tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.

Job Summary

We are seeking a highly technical and execution‑focused Head of Platform Product Reliability to lead reliability engineering across Etched's server, rack and datacenter platform products.

This role owns system‑level product reliability from architecture through fleet deployment. You will define reliability strategy, qualification methodologies, accelerated stress testing programs, failure analysis processes and long‑term reliability standards for complex AI infrastructure systems. This team focuses specifically on product reliability engineering for platform hardware and deployed systems – ensuring every Etched product ships with the reliability profile that enterprise and hyperscale customers demand.

You will work cross‑functionally with Platform Engineering, Mechanical Engineering, Thermal, Firmware, Manufacturing, Supply Chain, Datacenter Operations and Program teams to ensure Etched products achieve exceptional reliability at scale.

Key Responsibilities
  • Define and own the end‑to‑end reliability strategy for AI servers, accelerator platforms, rack systems and datacenter infrastructure, from design requirements through field deployment
  • Establish reliability requirements, qualification standards and validation methodologies that scale across product generations
  • Build and institutionalise reliability engineering processes spanning the full product life‑cycle:
    • EVT / DVT / PVT qualification gates and exit criteria
    • Accelerated life testing (ALT) and accelerated stress testing (AST)
    • Environmental testing: temperature, humidity, altitude, contamination
    • HALT / HASS programmes for design margin and production screening
    • Vibration, shock and transportation stress testing
    • Power cycling, thermal cycling and long‑duration soak testing
  • Lead root‑cause investigations for reliability failures surfaced during development, manufacturing and field deployment, driving corrective actions across hardware, firmware, thermal and mechanical domains
  • Develop comprehensive system reliability models including MTBF projections, FIT rate analysis, Weibull lifetime modelling, component derating methodologies and reliability growth tracking
  • Ensure reliability is considered early, partnering with Platform Engineering architects and design leads so reliability requirements shape decisions before they become expensive to change
  • Work closely with ODMs, JDMs, contract manufacturers and component suppliers to validate and enforce long‑term platform reliability commitments
  • Build fleet reliability infrastructure: telemetry analysis pipelines, field feedback loops and monitoring frameworks that give Etched visibility into deployed system health at scale
  • Drive reliability sign‑off criteria and lead product release readiness reviews across engineering and programme teams
  • Build and lead a high‑performing product reliability engineering organisation – hiring, developing and retaining technical talent as the company scales
You may be a good fit if you have (Must‑have qualifications)
  • BS, MS or PhD in Electrical Engineering, Mechanical Engineering, Reliability Engineering or a related technical field
  • 10+ years of reliability engineering experience in hardware‑centric organisations, with meaningful time spent on complex systems rather than component‑level work
  • Experience leading reliability programmes for one or more of:
    • AI accelerator or GPU‑class compute systems
    • Hyperscale or cloud server infrastructure
    • Networking platforms, storage systems or rack‑scale infrastructure
  • Deep understanding of system‑level failure mechanisms – including thermal, power delivery, mechanical and connector/interconnect failure modes – and how design decisions affect long‑term field reliability
  • Hands‑on experience with FMEA, Weibull analysis, HALT/HASS, qualification planning, failure analysis methodologies and reliability statistics and modelling
  • A track record of driving cross‑functional root‑cause investigations in fast‑moving hardware organisations where schedule pressure is real and accountability is high
  • Strong technical judgment – capable of making defensible trade‑offs between reliability targets, cost, schedule and performance without losing sight of customer expectations
  • Excellent communication skills and the credibility to influence design decisions with engineering leads, programme managers and executive stakeholders
Strong candidates may also have experience with (Nice‑to‑have qualifications)
  • Experience with liquid‑cooled systems, high‑density power delivery or thermal management for high‑power AI infrastructure
  • Direct experience supporting hyperscale or cloud datacenter deployments at scale, including customer‑facing reliability commitments and SLA management
  • Demonstrated experience building a reliability organisation from early‑stage – establishing processes, tooling and team norms in environments without established infrastructure
  • Familiarity with fleet telemetry systems, large‑scale field reliability analytics and data‑driven approaches to proactive reliability management
  • Experience working closely with ODM or JDM partners in Taiwan or broader Asia, including NPI support and on‑site qualification engagement
  • Background in high‑speed digital systems, GPU compute platforms or accelerator‑based architectures – with an understanding of how these affect system‑level reliability behaviour
Benefits
  • Medical, dental and vision packages with generous premium coverage
    • $500 per month credit for waiving medical benefits
  • Housing subsidy of $2k per month for those living within walking distance of the office
  • Relocation support for those moving to San Jose (Santana Row)
  • Various wellness benefits covering fitness, mental health and more
  • Daily lunch and dinner in our office
  • Unlimited compute budget subject to ROI justification
How we’re different

Etched believes in the Bitter Lesson. We are the first inference‑focused frontier AI system, betting early on transformer and transformer‑like architectures and on increasing model sizes. Our addressable market is the entirety of inference, unlike many of our competitors.

We are a fully in‑person team in San Jose (Santana Row) and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all our technical staff to contribute to both and work across disciplines as needed.

#J-18808-Ljbffr