1

Reliability Engineer Manager Jobs in Seattle, WA

Sr. SRE Consultant

Seattle, WA · On-site

$64.75 - $86.25/hr

Role: Sr. SRE (Very Strong Technical SRE) Location: Seattle, WA WFO: Mandatory (3 days/week) Short ... Experience in ITSM process including Incident, Problem, and Change management. * Experience in ...

The Apple Service Engineering - Data Streaming SRE team is looking for Site Reliability Engineers with experience developing processes, tools, and automation for managing distributed systems in ...

You will utilize yourstrongbackground in deploying, managing, and maintainingproduction systems ... As a Sr. Site Reliability Engineer, youwilltake independent responsibility for building and ...

New

The Apple Service Engineering - Data Streaming SRE team is looking for Site Reliability Engineers with experience developing processes, tools, and automation for managing distributed systems in ...

Standardize workflows for change management, deployment, and incident response, replacing manual ... Reliability Engineering, DevOps, or IT Operations within complex environments* Demonstrated ...

Senior Site Reliability Engineer

Bellevue, WA

$64.25 - $85.50/hr

The Senior Site Reliability Engineer Opportunity Reporting to the Manager, Site Reliability Engineering , this role will help build, improve, and maintain our cloud platform services by designing and ...

Showing results 21-40

Reliability Engineer Manager information

See Seattle, WA salary details

$69.4K

$134.3K

$160.5K

How much do reliability engineer manager jobs pay per year?

As of Aug 11, 2026, the average yearly pay for reliability engineer manager in Seattle, WA is $134,256.00, according to ZipRecruiter salary data. Most workers in this role earn between $116,600.00 and $146,800.00 per year, depending on experience, location, and employer.

What does a reliability engineer manager do?

A Reliability Engineer Manager oversees teams responsible for improving the reliability and performance of systems, machinery, or processes within an organization. They develop maintenance strategies, lead root cause analyses of failures, and implement best practices to minimize downtime and costs. Additionally, they collaborate with other departments to ensure that reliability goals align with business objectives and compliance standards. Their role is crucial in industries such as manufacturing, energy, and technology, where system uptime and safety are critical.

What are some common challenges reliability engineer managers face when balancing long-term reliability improvements with immediate operational demands?

Reliability Engineer Managers often need to prioritize urgent maintenance issues while also driving long-term reliability initiatives. Balancing these competing demands can be challenging, as immediate equipment failures may require quick fixes that temporarily interrupt ongoing improvement projects. Effective managers work closely with operations, maintenance, and engineering teams to communicate priorities, allocate resources, and implement sustainable solutions that address root causes rather than just symptoms. This role typically involves using data-driven decision-making and fostering a culture of proactive maintenance and continuous improvement.

What are the key skills and qualifications needed to thrive as a reliability engineer manager?

To thrive as a Reliability Engineer Manager, you need a strong background in engineering principles, reliability analysis, and maintenance strategies, typically supported by a degree in engineering and experience in reliability roles. Familiarity with reliability-centered maintenance (RCM), failure mode and effects analysis (FMEA), and asset management software such as SAP or Maximo is common, along with certifications like Certified Reliability Engineer (CRE). Leadership, problem-solving, and effective communication are vital soft skills for managing teams and driving cross-functional initiatives. These competencies are crucial for minimizing downtime, optimizing equipment performance, and ensuring long-term operational efficiency.

What is the difference between Reliability Engineer Manager vs Reliability Engineer?

AspectReliability EngineerReliability Engineer Manager
Required CredentialsBachelor's in Engineering or related field; certifications like CRC, CRESame as Reliability Engineer, plus leadership experience
Work EnvironmentDesign, analyze, and improve system reliability; often in teamsOversees Reliability Engineers; manages projects and teams
Employer & Industry UsageManufacturing, aerospace, energy, automotiveSame industries, with added managerial responsibilities
Common Search & ComparisonFocuses on technical skills and hands-on reliability tasksFocuses on leadership, team management, and strategic planning

The main difference between a Reliability Engineer and a Reliability Engineer Manager lies in their responsibilities. The Reliability Engineer focuses on technical analysis and system improvements, while the Reliability Engineer Manager oversees teams, manages projects, and develops strategies to enhance reliability across the organization.

What are the most commonly searched types of Reliability Engineer jobs in Seattle, WA? The most popular types of Reliability Engineer jobs in Seattle, WA are:
What are popular job titles related to Reliability Engineer Manager jobs in Seattle, WA? For Reliability Engineer Manager jobs in Seattle, WA, the most frequently searched job titles are:
What job categories do people searching Reliability Engineer Manager jobs in Seattle, WA look for? The top searched job categories for Reliability Engineer Manager jobs in Seattle, WA are:
What cities near Seattle, WA are hiring for Reliability Engineer Manager jobs? Cities near Seattle, WA with the most Reliability Engineer Manager job openings:

SRE - Senior AI Platform Reliability Engineer

HTC Global Services

Seattle, WA • On-site

$65 - $86.25/hr

Other

Medical, Dental, Vision, Life, Retirement, PTO

Re-posted 3 days ago


Job description

Are you a Senior Site Reliability Engineer with direct experience operating AI or machine learning platforms in large-scale production environments? We are looking for you to join our growing team working a hybrid schedule at one of our three locations either in Seattle, Burbak or Orlando. If working on a highly collaborative cutting edge technology team and in a job that is not a traditional infrastructure-only SRE role then this is the position for you!

The selected engineer will help build, scale, and operate the cloud and Kubernetes infrastructure that enables enterprise AI capabilities across the company and the selected engineer will help design, scale, and operate the cloud and Kubernetes infrastructure supporting enterprise AI workloads.

You do not need to be an AI model developers, data scientists, or LLM experts. instead your day to day focus and experience with the infrastructure and operational challenges involved in:

  • Deploying AI models or AI services into production
  • Operating AI platforms at enterprise scale
  • Supporting AI workloads across cloud environments
  • Scaling distributed AI services
  • Building reliable infrastructure for model inference, APIs, and AI-enabled applications
  • Monitoring the performance, availability, and capacity of production AI platforms

The ideal candidate combines strong AI platform infrastructure experience with deep expertise in SRE, Kubernetes, multi-cloud engineering, Infrastructure as Code, observability, and production reliability.

Key Responsibilities
  • Lead the design, implementation, and operation of highly available infrastructure supporting enterprise AI platforms and services.
  • Build and operate Kubernetes-based environments used to deploy and scale AI workloads.
  • Support the production deployment of AI models, inference services, AI APIs, agents, and related platform capabilities.
  • Design infrastructure that enables AI workloads across distributed and multi-cloud environments.
  • Partner with AI engineers, platform engineers, architects, and application teams to move AI services from development into reliable production environments.
  • Establish deployment, scaling, capacity, and reliability patterns for AI-powered services.
  • Design and maintain Kubernetes infrastructure using Helm and Terraform.
  • Support AI platform dependencies such as model gateways, vector or operational data stores, messaging platforms, secrets management, and API services.
  • Develop scalable platform solutions capable of maintaining 99.99% availability.
  • Lead capacity planning for AI workloads, including compute, memory, storage, network, and service dependencies.
  • Build automated CI/CD pipelines using Harness or comparable enterprise deployment platforms.
  • Implement blue/green deployments, canary releases, automated rollback, and feature-flag strategies.
  • Build observability for AI platforms using metrics, logs, traces, service-level indicators, and workload-specific health signals.
  • Troubleshoot complex production issues involving AI services, Kubernetes, cloud infrastructure, databases, messaging, networking, and distributed systems.
  • Lead incident response, root-cause analysis, and permanent corrective-action planning.
  • Mentor SRE, DevOps, and platform engineers and establish engineering and operational standards.
  • Ensure AI infrastructure aligns with enterprise security, governance, compliance, and resiliency requirements.
Required Qualifications
  • 7+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, cloud infrastructure, or a related field.
  • Direct experience supporting AI, machine learning, or model-serving platforms in production.
  • Experience deploying or operating AI models, inference services, AI APIs, agents, or AI-enabled applications in cloud environments.
  • Experience supporting AI workloads at enterprise or high-traffic scale.
  • Strong understanding of the infrastructure required to move AI services from development into production.
  • Expert-level Kubernetes administration and production operations experience.
  • Strong Helm and Terraform experience.
  • Experience designing scalable, highly available, distributed cloud infrastructure.
  • Hands-on experience with Google Cloud Platform, with additional AWS or Azure experience.
  • Experience building automated deployment pipelines using Harness or a comparable enterprise CI/CD platform.
  • Strong scripting and automation skills using Python, Bash, and YAML.
  • Experience supporting production databases, messaging, and platform services such as PostgreSQL, Redis, Kafka, MongoDB, and Vault.
  • Experience implementing observability using OpenTelemetry, Prometheus, Grafana, Splunk, AppDynamics, or similar tools.
  • Strong production troubleshooting and incident-management experience across cloud-native distributed systems.
  • Experience with capacity planning, performance tuning, disaster recovery, and high-availability design.
  • Strong communication and technical leadership skills.
Preferred Qualifications
  • Experience supporting generative AI, LLM, model inference, or AI-agent platforms.
  • Experience operating model gateways, AI APIs, GPU-enabled workloads, or distributed inference services.
  • Experience with AI platform technologies such as LiteLLM, Open WebUI, model-serving frameworks, vector databases, or similar platforms.
  • Experience defining SLOs, SLIs, error budgets, and reliability standards for AI services.
  • Experience with Google Cloud Platform-based AI infrastructure and services.
  • Experience supporting high-volume or globally distributed AI platforms.
  • Experience with progressive delivery, chaos engineering, and automated resilience testing.
  • Cloud or Kubernetes certifications.

What Makes HTC A Great Place To Build Your Future

HTC Global Services wants you to join our team. Come build new things with us and advance your career. At HTC Global, you ll collaborate with experts, work alongside clients, and be part of high-performing teams driving success together. You ll have long-term opportunities to grow your career and develop skills in the latest emerging technologies.

At HTC Global Services, our employees have access to a comprehensive benefits package. Benefits can include Group Health (Medical, Dental, and Vision), Paid Time Off, Paid Holidays, 401(k) matching, Group Life and Disability insurance, Professional Development opportunities, Wellness programs, and a variety of other perks.

Our success as a company is built on inclusion and diversity. HTC Global Services is committed to providing a workplace free from discrimination and harassment, where every employee is treated with dignity and respect. We celebrate differences and believe that diverse cultures, perspectives, and skills drive innovation and success. HTC is an Equal Opportunity Employer and a proud National Minority Supplier. We seek to empower each individual, fostering an environment where everyone feels valued, included, and respected.

#LI-NC1 #LI-DT1 #LI-Remote #Hiring #AIJobs #SREJobs #SeattleJobs #BurbankJobs #OrlandoJobs