1

Site Reliability Engineer Manager Jobs in Oregon

Senior Engineering Manager, Site Reliability

OR · On-site +1

$57 - $75.75/hr

You will serve as the accountable leader of the SRE function, translating reliability strategy into ... Manage and develop a team focused on incident management, observability, operational readiness, and ...

Site Reliability Engineer II

OR · On-site +1

$57 - $75.75/hr

The Site Reliability Engineering (SRE) team integrates software and systems engineering to design and manage large-scale, distributed, and fault-tolerant systems. The team is responsible for ensuring ...

Senior Software Engineer, Site Reliability

OR · On-site +1

$57 - $75.75/hr

The Team Upstart's Site Reliability Engineering (SRE) team owns the reliability, resiliency, and ... Experience with on-call and incident management environments * Experience with observability ...

Site Reliability Engineer

OR · On-site +1

$57 - $75.75/hr

Additionally, the SRE I role will be an expert in Infrastructure as Code (IaC), AWS Cloud ... Design, implement, manage, and optimize robust CI/CD pipelines using tools like GitHub Actions and ...

Site Reliability Engineer TELCOR Inc, a leading innovator in laboratory software, is looking for a ... environments, as well as manage production infrastructure and deployment workflows across ...

Site Reliability Engineer

Lake Oswego, OR · On-site

$105K - $145K/yr

Strong proficiency in PowerShell and Bash, along with experience managing GitHub Actions and Azure ... , SRE, Azure, Terraform, Site Reliability Engineer, Enterprise Software, CI/CD, New Relic ...

Software Engineer, Site Reliability

OR · On-site +1

$57 - $75.75/hr

The Team Upstart's Site Reliability Engineering team enables engineers to operate reliable ... Experience building internal reliability, observability, incident management, or operational ...

New

Sr. Site Reliability Engineer

OR · On-site +1

$57 - $75.75/hr

FreedomPay is seeking an experienced Senior Site Reliability Engineer to help ensure the highest ... Experience implementing enterprise incident management practices. * Experience building AIOps or ML ...

Senior Site Reliability Engineer- Remote

OR · On-site +1

$57 - $75.75/hr

... central Site Reliability Engineering team. You will be responsible for building and leading ... You will also own the areas of incident management and response, post-mortem analysis including ...

Site Reliability Engineer (Tu-Sat, night shift)

OR · On-site +1

$57 - $75.75/hr

The ideal candidate has prior experience managing cloud-based SaaS applications and strives to ... Site Reliability Engineering (SRE) is a growing team that partners closely with Product Engineering ...

Senior Site Reliability Engineer - Compute Platforms

OR · On-site +1

$57 - $75.75/hr

This is a deeply technical role requiring expert-level understanding of compute hardware management ... You will also collaborate with platform and SRE teams to maintain secure, performant, and multi ...

Senior Staff Site Reliability Engineer

OR · On-site +1

$170K - $227K/yr

As a Ping Identity Senior Staff Site Reliability Engineer, you will be involved in every facet of ... You have experience with instrumentation and management of automated deployments * You have ...

next page

Showing results 1-20

Site Reliability Engineer Manager information

See Oregon salary details

$11

$67

$97

How much do site reliability engineer manager jobs pay per hour?

As of Aug 24, 2026, the average hourly pay for site reliability engineer manager in Oregon is $67.39, according to ZipRecruiter salary data. Most workers in this role earn between $57.93 and $77.02 per hour, depending on experience, location, and employer.

What is a site reliability engineer manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What are the key skills and qualifications needed to thrive as a site reliability engineer manager?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.

How does a site reliability engineer manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How much do site reliability engineer managers get paid?

Site Reliability Engineer Managers typically earn between $120,000 and $180,000 annually, depending on experience, location, and company size. They often oversee teams responsible for system reliability, incident response, and infrastructure automation, requiring strong leadership and technical skills.

Is a Site Reliability Engineer Manager a stressful job?

A Site Reliability Engineer Manager role can be stressful due to the responsibility of maintaining system uptime, managing incident responses, and ensuring reliability across complex infrastructure. The job often involves working under pressure, handling outages, and coordinating teams, but it also offers opportunities for problem-solving and leadership. Stress levels vary depending on company size, team structure, and workload management skills.

What are the most commonly searched types of Site Reliability Engineer jobs in Oregon?

The most popular types of Site Reliability Engineer jobs in Oregon are:

What cities in Oregon are hiring for Site Reliability Engineer Manager jobs?

Cities in Oregon with the most Site Reliability Engineer Manager job openings:

Infographic showing various Site Reliability Engineer Manager job openings in Oregon as of August 2026, with employment types broken down into 84% Full Time, 15% Part Time, and 1% Contract. Highlights an 79% Physical, 3% Hybrid, and 18% Remote job distribution, with an average salary of $140,178 per year, or $67.4 per hour.

Senior Engineering Manager, Site Reliability

OR • On-site, Remote


Upstart

7.6

Company rating: 7.6 out of 10

Based on 6 frontline employees who took The Breakroom Quiz

Good employer

Respectful managers

Good training


$57 - $75.75/hr

Full-time

Re-posted 7 days ago


Job description

The Team: 

The Site Reliability Engineering (SRE) team enables Upstart's engineering organization to operate reliable, observable, and resilient systems at scale. The team owns company-wide incident management practices, reliability standards, operational readiness, and the capabilities that help engineering teams identify, respond to, and learn from production issues.

Our goal is to make reliability an integrated part of how software is designed, delivered, and operated. We are building a model where engineering teams have the trusted signals, automated safeguards, and operational practices needed to move quickly while protecting our customers and business.

The team advances observability, incident detection and response, service level objectives, operational readiness, and systemic improvements based on incident learnings. SRE partners across product engineering, infrastructure, security, and platform teams to improve reliability at scale.

The Role:

As the Senior Engineering Manager of Site Reliability Engineering, you will lead a team responsible for improving the reliability and operational maturity of Upstart's products and services. You will drive high impact improvements across incident management, observability, operational readiness, and reliability engineering.

You will serve as the accountable leader of the SRE function, translating reliability strategy into focused plans, clear ownership, and measurable outcomes. You will partner closely with engineering leaders to establish reliability expectations, identify systemic risks, and build scalable capabilities that enable teams to operate services safely and independently.

This role is suited for a hands-on leader with strong technical judgment, disciplined execution, and a track record of building high performing teams. You will balance immediate operational needs with durable improvements that reduce risk, strengthen resilience, and improve how Upstart learns from production.

How you'll make an impact

Team Leadership and Execution
  • Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering
  • Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function
  • Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures
  • Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts
  • Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership
  • Set a high bar for technical quality, operating rigor, and executive communication
  • Develop engineers and leaders who can independently own complex reliability initiatives
Incident Management and Learning
  • Evolve Upstart's incident management program to improve detection, response, coordination, communication, and recovery
  • Establish clear standards for managing high severity incidents and provide visible leadership during critical events
  • Improve postmortem quality and ensure incident learnings result in durable engineering improvements
  • Identify recurring failure patterns and drive systemic solutions across teams
  • Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction
Observability and Reliability Engineering
  • Improve the quality, accessibility, and trustworthiness of signals used to understand production health
  • Drive consistent practices across metrics, logs, traces, alerting, and service health
  • Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions
  • Reduce detection gaps, noisy alerts, manual investigation, and recurring operational toil
  • Define measurable reliability outcomes and use data to prioritize investments and communicate impact
  • Partner with platform and product engineering teams to embed reliability into standard engineering workflows
Operational Readiness and Resilience
  • Establish scalable operational readiness standards for new services, major launches, and architectural changes
  • Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response
  • Identify systemic reliability risks and partner with engineering teams to prioritize and address them
  • Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards
  • Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through
  • Align stakeholders and dependencies before critical launches and engineering decisions

Minimum Qualifications 

  • 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering
  • Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems
  • Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes
  • Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations
  • Experience leading high severity incident response and improving incident management practices at scale
  • Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes
  • Track record of hiring, developing, and retaining high performing engineers and engineering leaders
  • Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams

Preferred Qualifications

  • Experience operating large scale, highly available distributed systems
  • Experience implementing or evolving service-level objectives and error-budget practices
  • Experience with observability platforms such as Datadog, Grafana, Prometheus, OpenTelemetry, or similar technologies
  • Experience developing incident management, operational readiness, or resilience programs across a large engineering organization
  • Familiarity with Kubernetes, AWS, and modern cloud native architectures
  • Experience supporting major platform or architectural transitions
  • Strong product mindset when building internal reliability capabilities
  • Experience establishing executive level reliability reporting and operating reviews

Position location This role is available in the following locations: Remote

Travel requirements As a digital first company, the majority of your work can be accomplished remotely. The majority of our employees can live and work anywhere in the U.S but are encouraged to to still spend high quality time in-person collaborating via regular onsites. The in-person sessions' cadence varies depending on the team and role; most teams meet once or twice per quarter for 2-4 consecutive days at a time.

#LI-REMOTE

#LI-MidSenior


What Upstart employees say

Pay

Hours and flexibility

Workplace

Get the full story on Breakroom