1

Reliability Engineer Jobs in Redmond, WA (NOW HIRING)

We are looking for an experienced SRE leader with the skills and passion to make a significant impact on our ecosystem. You will have a wide array of projects to tackle, with ample opportunities for ...

Sr. SRE Consultant

Seattle, WA · On-site

$64.75 - $86.25/hr

Role: Sr. SRE (Very Strong Technical SRE) Location: Seattle, WA WFO: Mandatory (3 days/week) Short JD: * Job Summary/Role Description (General Information) * As a Senior Site Reliability Engineer ...

Principal Site Reliability Engineer

Bellevue, WA · On-site

$64.25 - $85.50/hr

We are looking for an experienced SRE leader with the skills and passion to make a significant impact on our ecosystem. You will have a wide array of projects to tackle, with ample opportunities for ...

Principal Site Reliability Engineer

Bellevue, WA · On-site

$64.25 - $85.50/hr

We are looking for an experienced SRE leader with the skills and passion to make a significant impact on our ecosystem. You will have a wide array of projects to tackle, with ample opportunities for ...

Sr. Site Reliability Engineer

Seattle, WA · On-site

$175 - $200/hr

Asa Sr. Site Reliability Engineer (SRE) inPitchBook'sengineeringdivision, you will becreating and evolving systems to automatically runour suite of products and servicesreliablyand consistently.As ...

Our SRE team combines software engineering, systems engineering, and Devops practices to build and run large-scale, massively distributed, fault-tolerant systems. Our software ensures that Apple ...

Site Reliability Engineer IV

Mountlake Terrace, WA · On-site

$135.60 - $230.50/hr

## Site Reliability Engineer IVApplylocations: Mountlake Terrace WAtime type: Full timeposted on: Posted 2 Days Agojob requisition id: R28853**Workforce Classification:**Hybrid**Join Our Team: Do ...

Our SRE team combines software engineering, systems engineering, and Devops practices to build and run large-scale, massively distributed, fault-tolerant systems. Our software ensures that Apple ...

Senior Site Reliability Engineer

Bellevue, WA · On-site

$64.25 - $85.50/hr

The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering , this role will help build, improve, and maintain our cloud platform services by designing and ...

Showing results 21-40

Reliability Engineer information

See Redmond, WA salary details

$68.3K

$132.1K

$157.9K

How much do reliability engineer jobs pay per year?

As of Aug 13, 2026, the average yearly pay for reliability engineer in Redmond, WA is $132,124.00, according to ZipRecruiter salary data. Most workers in this role earn between $114,800.00 and $144,500.00 per year, depending on experience, location, and employer.

What are some typical challenges reliability engineers face when implementing preventive maintenance strategies?

Reliability Engineers often encounter challenges such as balancing preventive maintenance schedules with production demands, ensuring buy-in from operations teams, and accurately predicting equipment failures. They must analyze large sets of historical data to identify trends and root causes, which can be complex in facilities with diverse machinery. Collaboration with maintenance, operations, and engineering teams is essential to develop effective strategies that minimize downtime while optimizing resources.

How much do reliability engineers get paid?

Reliability engineers typically earn a median annual salary ranging from $70,000 to $110,000, depending on experience, location, and industry. Senior or specialized reliability engineers with certifications and advanced skills can earn higher salaries, often exceeding $120,000 annually.

What are the key skills and qualifications needed to thrive as a reliability engineer, and why are they important?

To thrive as a Reliability Engineer, you need a solid background in engineering principles, failure analysis, and reliability modeling, typically with a degree in engineering or a related field. Familiarity with tools such as FMEA, Root Cause Analysis (RCA), reliability-centered maintenance (RCM) software, and certifications like Certified Reliability Engineer (CRE) are highly valued. Strong problem-solving abilities, attention to detail, and effective communication are crucial soft skills in this role. These skills ensure systems are dependable, downtime is minimized, and organizational performance and safety are optimized.

What is the difference between Reliability Engineer vs Maintenance Engineer?

AspectReliability EngineerMaintenance Engineer
CredentialsTypically requires engineering degree, certifications in reliability or asset managementOften requires engineering or technical diploma, certifications in maintenance or equipment repair
Work EnvironmentFocuses on analysis, design, and improvement of systems for reliabilityHands-on maintenance, repair, and troubleshooting of equipment
Industry UsageCommon in manufacturing, energy, aerospace, and industrial sectorsPrevalent in manufacturing, facilities, and industrial plants

Reliability Engineers focus on designing and improving systems to prevent failures, using data analysis and modeling. Maintenance Engineers perform hands-on repairs and upkeep of equipment to ensure operational continuity. While both roles aim to optimize equipment performance, Reliability Engineers work proactively on system reliability, whereas Maintenance Engineers handle reactive and scheduled maintenance tasks.

What does a reliability engineer do?

As a reliability engineer, your duties are to test and evaluate the manufacturing of products and components and ensure that the procedures are efficient and do not lead to abnormally high maintenance or operational costs. Your other responsibilities are to find solutions to product reliability risks. You may manage risk in a supply chain, develop loss prevention strategies, and track the entire lifecycle of product development, from building prototypes to moving a product into full-scale production. You analyze information from department heads and recommend strategies to reduce risk and ensure that the product works reliably.

What is a reliability engineer?

Reliability Engineers are professionals responsible for ensuring that systems, equipment, or processes function consistently and efficiently over time. They analyze data, identify potential points of failure, and develop maintenance strategies to improve system reliability and minimize downtime. Their work spans various industries, including manufacturing, energy, and technology, and often involves collaborating with design, operations, and maintenance teams. By implementing reliability-centered maintenance and predictive analysis, they help organizations save costs and increase safety.

What are popular job titles related to Reliability Engineer jobs in Redmond, WA?

For Reliability Engineer jobs in Redmond, WA, the most frequently searched job titles are:

What job categories do people searching Reliability Engineer jobs in Redmond, WA look for?

The top searched job categories for Reliability Engineer jobs in Redmond, WA are:

What cities near Redmond, WA are hiring for Reliability Engineer jobs?

Cities near Redmond, WA with the most Reliability Engineer job openings:

Infographic showing various Reliability Engineer job openings in Redmond, WA as of August 2026, with employment types broken down into 82% Full Time, and 18% Contract. Highlights an 91% In-person, and 9% Remote job distribution, with an average salary of $132,124 per year, or $63.5 per hour.

Principal Site Reliability Engineer

iSpot

Bellevue, WA

$64.25 - $85.50/hr

Full-time

Re-posted 24 days ago


Job description

What You'll Be Part Of:

iSpot.tv is changing how brands, agencies, and networks measure and assess the impact of TV advertising. We deal with BIG data, operating mainly in AWS with multiple Kubernetes clusters and thousands of servers. We are looking for an experienced SRE leader with the skills and passion to make a significant impact on our ecosystem. You will have a wide array of projects to tackle, with ample opportunities for growth.

You will be a key member of our SRE leadership team, focused on empowering developers to build, test, and deploy applications faster and more efficiently. You will both lead the team and remain hands-on in designing, building, and maintaining the tools, platforms, and processes that improve our engineering teams' productivity and streamline the software development lifecycle. Your work will directly impact developer happiness and the speed at which we can deliver innovative features to our customers.

Responsibilities:

We are seeking a seasoned and strategic Lead/Principal Site Reliability Engineer to drive the reliability, scalability, and performance of our core production systems while significantly enhancing the internal developer experience. This role sits at the intersection of operations and development, requiring deep technical expertise, strong leadership, and a passion for optimizing the entire software development lifecycle (SDLC).

Our team consists of senior engineers who work together with minimal supervision to attain those goals. Candidates must possess deep operational experience with AWS and Kubernetes to support teams utilizing these systems. You will lead the technical direction of the team while remaining a key individual contributor. You will be responsible for creating a culture of engineering excellence, designing self-service platforms, and fostering alignment across all engineering teams to accelerate product delivery and maintain world-class service stability.The key responsibilities are:

  1. System Reliability and Operations (SRE Focus)
  • Platform Design and Management: Architect, build, and maintain scalable, highly available, and reliable cloud infrastructure in AWS leveraging modern container orchestration technologies.
  • Data Pipeline Reliability: Serve as the reliability and cost optimization expert for high-volume, data-intensive workloads. Focus on optimizing and ensuring the stability of distributed data processing engines, specifically Apache Spark and related ecosystems (e.g., EMR, Databricks, Glue).
  • Observability and Monitoring: Establish comprehensive observability practices by defining SLIs/SLOs, implementing advanced monitoring, alerting, and logging solutions to quickly identify and resolve system anomalies.
  • Automation: Drive automation across all operational aspects, including infrastructure provisioning (Terraform), scaling, deployment, and incident response, minimizing toil and manual effort.
  • Incident Management: Lead and participate in the incident response lifecycle, performing thorough post-mortems to derive actionable insights and implement preventative measures to improve system resilience.
  • AIOps: Define and champion the strategic roadmap for AI/ML integration within SRE, establishing organizational best practices for AIOps, automated incident remediation, Toil Reduction via LLMs, and Automated Root Cause Analysis (RCA) and the governance of LLM-driven tooling to enhance system observability and resilience.
  1. Developer Experience and Productivity (DevEx Focus)
  • Platform Strategy: Design, implement, and champion self-service tools, internal developer portals, and services that empower engineering teams to manage their infrastructure and deployments independently and efficiently.
  • AI Developer Tools: Lead the standardization of AI developer assistants by architecting and maintaining global 'steering files' and context-configuration standards, ensuring AI-generated code aligns with our specific patterns, security protocols, and architectural guardrails.
  • CI/CD Optimization: Own and continuously improve the CI/CD pipelines, reducing build times, streamlining deployment workflows, and integrating best practices for testing, security (Shift Left), and code quality. Maintain and improve our container orchestration and deployment tools, leveraging Kubernetes, Helm, and ArgoCD to create seamless developer workflows.
  • KPIs: Develop, implement, and maintain a set of key performance indicators (KPIs) to measure and improve the developer experience across all of Engineering.
  • Mentorship and Documentation: Guide and mentor senior engineers, promoting SRE/DevEx principles. Develop clear, comprehensive documentation and tutorials to ensure seamless adoption of new tools and platforms.
  • Cost and Efficiency: Strategically identify and implement opportunities for cloud cost optimization and resource efficiency without compromising reliability or performance.

III. Strategic Leadership and Cross-Team Alignment

  • Architecting the Roadmap: Define, champion, and communicate the long-term technical roadmap for the SRE and DevEx platforms, balancing immediate operational needs with strategic, future-state goals.
  • Driving Cross-Team Alignment: Act as a critical liaison between infrastructure, security, and product development teams. Proactively drive cross-team alignment on architectural standards, tooling choices, and development workflows to ensure consistency and shared accountability for system health.
  • Bottleneck Identification and Mitigation: Systematically identify engineering bottlenecks, friction points, and points of organizational toil within the SDLC. Implement targeted solutions-whether technical, process-based, or organizational-to mitigate these constraints and enhance overall engineering velocity.
  • Planning and Execution: Collaborate with engineering leadership to transform the strategic roadmap into actionable, prioritized plans, securing cross-functional buy-in and resources for successful execution.

 Qualifications and Education Requirements:

  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 10+ years of relevant experience in software engineering, cloud architecture, and/or Site Reliability Engineering, with at least 3 years in a leadership or lead contributor role.
  • Deep expertise of AWS, including EKS, ECR, RDS, SQS/SNS, VPC, MWAA and S3.
  • Strong proficiency in Infrastructure as Code (IaC) tools (e.g., Terraform, CloudFormation).
  • Specialized experience in optimizing large-scale data platforms, specifically with Apache Spark. Proven ability to profile, troubleshoot, and tune Spark jobs for performance, cost, and reliability.
  • 5+ years of experience with Kubernetes and containerization in general, including associated tools (kubectl, Helm, ArgoCD).
  • Strong knowledge of AWS cost optimization.
  • TCP/IP networking, including routing and AWS security groups.
  • Excellent knowledge of CI/CD concepts and experience developing associated pipelines in CircleCI.
  • Proficient in high-level scripting languages, including shell scripting, Python, and/or JavaScript.
  • Experience with OTel and monitoring tools such as Splunk or DataDog. Experience with native AI observability tools is a plus.
  • Experience with evaluating and rolling out GenAI tools for improving developer efficiency.
  • Excellent communication, collaboration, and stakeholder management skills, with proven experience driving technical initiatives across multiple teams.
  • Experience with researching and selecting new/modern developer toolsets and assisting teams in adopting them including vendor assessments, security assessments and procurement process.
  • Experience in Ad-Tech or "BIG Data" processing organization is highly preferred

Target cash compensation range: $163,620 - $212,710 USD Annually

We are committed to providing competitive, market-informed compensation. The cash compensation above includes base salary, variable commission for employees in eligible roles, and annual bonus targets for eligible roles. In addition to cash compensation, all full time iSpotters are eligible to participate in iSpot's equity plan to receive stock options. Non-exempt roles will also be eligible for (pre-approved) overtime pay. Individual compensation packages are influenced by different factors unique to each candidate, including their skills, experience, qualifications and other job-related reasons.

For more information on total rewards package, go HERE

Hybrid & Flexible Workplace Policy

iSpot supports a hybrid and flexible workplace. Depending on location and work responsibilities, employees may be designated as full-time or part-time office-based or a fully remote employee. A hybrid work schedule indicates that you work in the office some days and work from home other days. The best hybrid workplaces allow for flexibility while also encouraging consistency. 

Those local or living in surrounding areas to one of our offices (Bellevue, WA or New York, NY) will work a hybrid schedule, coming into their local office 1-3 days a week. While those in a role, not office-based and located further away from our offices, will work a fully remote schedule. If you have questions regarding exact details of our hybrid & flexible workplace policy, please let your recruiter know and they will discuss with you further.

#LI-Hybrid