1

Sre Manager Jobs in Oregon (NOW HIRING)

Senior Engineering Manager, Site Reliability

OR · On-site +1

$57 - $75.75/hr

You will serve as the accountable leader of the SRE function, translating reliability strategy into ... Manage and develop a team focused on incident management, observability, operational readiness, and ...

Site Reliability Engineer II

OR · On-site +1

$57 - $75.75/hr

The Site Reliability Engineering (SRE) team integrates software and systems engineering to design and manage large-scale, distributed, and fault-tolerant systems. The team is responsible for ensuring ...

Strong knowledge of SRE best practices and incident management protocols * Deep experience using and/or configuring New Relic, Data Dog, SumoLogic or similar observability tools * Proficiency in ...

Senior Software Engineer, Site Reliability

OR · On-site +1

$57 - $75.75/hr

The Team Upstart's Site Reliability Engineering (SRE) team owns the reliability, resiliency, and ... Experience with on-call and incident management environments * Experience with observability ...

Site Reliability Engineer

OR · On-site +1

$57 - $75.75/hr

Additionally, the SRE I role will be an expert in Infrastructure as Code (IaC), AWS Cloud ... Design, implement, manage, and optimize robust CI/CD pipelines using tools like GitHub Actions and ...

Site Reliability Engineer TELCOR Inc, a leading innovator in laboratory software, is looking for a ... environments, as well as manage production infrastructure and deployment workflows across ...

Strong proficiency in PowerShell and Bash, along with experience managing GitHub Actions and Azure ... , SRE, Azure, Terraform, Site Reliability Engineer, Enterprise Software, CI/CD, New Relic ...

Sr. Site Reliability Engineer

OR · On-site +1

$57 - $75.75/hr

FreedomPay is seeking an experienced Senior Site Reliability Engineer to help ensure the highest ... Experience implementing enterprise incident management practices. * Experience building AIOps or ML ...

Senior Site Reliability Engineer- Remote

OR · On-site +1

$57 - $75.75/hr

... central Site Reliability Engineering team. You will be responsible for building and leading ... You will also own the areas of incident management and response, post-mortem analysis including ...

Site Reliability Engineer

Portland, OR · On-site

$180K - $250K/yr

... manage access. SaaS apps, infrastructure, internal tools, AI agents-if it has a concept of "who ... engineering -partner with product teams to ensure new features are built with reliability in mind ...

New

Site Reliability Engineer (Tu-Sat, night shift)

OR · On-site +1

$57 - $75.75/hr

The ideal candidate has prior experience managing cloud-based SaaS applications and strives to ... Site Reliability Engineering (SRE) is a growing team that partners closely with Product Engineering ...

Senior Site Reliability Engineer - Compute Platforms

OR · On-site +1

$57 - $75.75/hr

This is a deeply technical role requiring expert-level understanding of compute hardware management ... You will also collaborate with platform and SRE teams to maintain secure, performant, and multi ...

Senior Staff Site Reliability Engineer

OR · On-site +1

$170K - $227K/yr

As a Ping Identity Senior Staff Site Reliability Engineer, you will be involved in every facet of ... You have experience with instrumentation and management of automated deployments * You have ...

next page

Showing results 1-20

Sre Manager information

See Oregon salary details

$65.6K

$124.2K

$178.2K

How much do sre manager jobs pay per year?

As of Aug 19, 2026, the average yearly pay for sre manager in Oregon is $124,218.00, according to ZipRecruiter salary data. Most workers in this role earn between $99,900.00 and $148,000.00 per year, depending on experience, location, and employer.

What is an SRE manager?

SRE Managers are leaders responsible for overseeing Site Reliability Engineering (SRE) teams. They ensure the reliability, scalability, and performance of software systems by guiding engineers in implementing best practices, automation, and monitoring processes. SRE Managers collaborate closely with development and operations teams to balance feature development with system stability. Their role also includes mentoring SREs, managing incident response, and driving improvements in system reliability and operational efficiency.

What are some common challenges SRE managers face when leading Site Reliability Engineering teams?

SRE Managers often encounter challenges balancing reliability with rapid development, ensuring their teams have the right mix of software engineering and operations skills. They must also foster a culture of continuous improvement while managing on-call rotations and incident response without causing burnout. Additionally, collaborating effectively with development and product teams to set realistic service level objectives (SLOs) and drive adoption of SRE best practices can require strong communication and negotiation skills.

What are the key skills and qualifications needed to thrive as an SRE manager, and why are they important?

To thrive as an SRE Manager, you need a deep understanding of site reliability engineering principles, strong experience with systems architecture, and a background in computer science or a related field. Familiarity with tools such as Kubernetes, Prometheus, cloud platforms, and CI/CD pipelines, as well as certifications like AWS Certified Solutions Architect, are commonly expected. Leadership, effective communication, and problem-solving skills are crucial for driving team performance and collaborating across departments. These skills ensure high system reliability, efficient incident management, and a culture of continuous improvement within technical organizations.

What is the difference between Sre Manager vs DevOps Engineer?

AspectSre ManagerDevOps Engineer
CredentialsTypically requires a Bachelor's/Master's in CS or related field, with certifications like AWS, Google Cloud, or KubernetesSimilar credentials, often with cloud certifications and scripting skills
Work EnvironmentLeads teams, manages incident response, and oversees reliability strategiesFocuses on automation, CI/CD pipelines, and infrastructure deployment
Industry UsageCommon in large tech companies, financial services, and cloud providersWidely used across startups, tech firms, and enterprises adopting DevOps practices

The Sre Manager and DevOps Engineer roles share overlapping skills in cloud computing, automation, and infrastructure management. While the Sre Manager oversees reliability and team coordination, the DevOps Engineer focuses on implementing automation tools and deployment pipelines. Both roles are crucial for modern IT operations, but the Sre Manager typically has a broader leadership responsibility, whereas the DevOps Engineer is more hands-on with technical implementation.

What does an SRE manager do?

An SRE manager oversees the reliability and performance of software systems, leading teams that implement automation, monitoring, and incident response processes. They coordinate efforts to ensure system availability, scalability, and efficiency, often using tools like monitoring dashboards and incident management platforms. Strong leadership, technical expertise, and understanding of service level objectives are essential in this role.

What are the most commonly searched types of Sre jobs in Oregon?

The most popular types of Sre jobs in Oregon are:

Infographic showing various Sre Manager job openings in Oregon as of August 2026, with employment types broken down into 43% Full Time, and 57% Contract. Highlights an 74% In-person, and 26% Remote job distribution, with an average salary of $124,218 per year, or $59.7 per hour.

Senior Engineering Manager, Site Reliability

Upstart

OR • On-site, Remote

$57 - $75.75/hr

Full-time

Re-posted 2 days ago


Upstart rating

7.6

Company rating: 7.6 out of 10

Based on 6 frontline employees who took The Breakroom Quiz


Job description

The Team: 

The Site Reliability Engineering (SRE) team enables Upstart's engineering organization to operate reliable, observable, and resilient systems at scale. The team owns company-wide incident management practices, reliability standards, operational readiness, and the capabilities that help engineering teams identify, respond to, and learn from production issues.

Our goal is to make reliability an integrated part of how software is designed, delivered, and operated. We are building a model where engineering teams have the trusted signals, automated safeguards, and operational practices needed to move quickly while protecting our customers and business.

The team advances observability, incident detection and response, service level objectives, operational readiness, and systemic improvements based on incident learnings. SRE partners across product engineering, infrastructure, security, and platform teams to improve reliability at scale.

The Role:

As the Senior Engineering Manager of Site Reliability Engineering, you will lead a team responsible for improving the reliability and operational maturity of Upstart's products and services. You will drive high impact improvements across incident management, observability, operational readiness, and reliability engineering.

You will serve as the accountable leader of the SRE function, translating reliability strategy into focused plans, clear ownership, and measurable outcomes. You will partner closely with engineering leaders to establish reliability expectations, identify systemic risks, and build scalable capabilities that enable teams to operate services safely and independently.

This role is suited for a hands-on leader with strong technical judgment, disciplined execution, and a track record of building high performing teams. You will balance immediate operational needs with durable improvements that reduce risk, strengthen resilience, and improve how Upstart learns from production.

How you'll make an impact

Team Leadership and Execution
  • Manage and develop a team focused on incident management, observability, operational readiness, and reliability engineering
  • Define a clear charter, priorities, roadmap, and measurable outcomes for the SRE function
  • Translate strategy into capacity aware plans with explicit trade offs, ownership, milestones, and success measures
  • Maintain visibility into delivery health, operational risks, and team performance, intervening early when execution drifts
  • Build a resilient operating model through cross-training, shared context, effective delegation, and clear primary and secondary ownership
  • Set a high bar for technical quality, operating rigor, and executive communication
  • Develop engineers and leaders who can independently own complex reliability initiatives
Incident Management and Learning
  • Evolve Upstart's incident management program to improve detection, response, coordination, communication, and recovery
  • Establish clear standards for managing high severity incidents and provide visible leadership during critical events
  • Improve postmortem quality and ensure incident learnings result in durable engineering improvements
  • Identify recurring failure patterns and drive systemic solutions across teams
  • Create strong feedback loops from incidents into roadmaps, service standards, operational readiness requirements, and measurable risk reduction
Observability and Reliability Engineering
  • Improve the quality, accessibility, and trustworthiness of signals used to understand production health
  • Drive consistent practices across metrics, logs, traces, alerting, and service health
  • Advance the use of service level objectives and customer impact signals to guide priorities and operational decisions
  • Reduce detection gaps, noisy alerts, manual investigation, and recurring operational toil
  • Define measurable reliability outcomes and use data to prioritize investments and communicate impact
  • Partner with platform and product engineering teams to embed reliability into standard engineering workflows
Operational Readiness and Resilience
  • Establish scalable operational readiness standards for new services, major launches, and architectural changes
  • Set clear expectations for service ownership, monitoring, capacity, failure handling, and incident response
  • Identify systemic reliability risks and partner with engineering teams to prioritize and address them
  • Improve resilience through automation, failure testing, recovery capabilities, and operational safeguards
  • Build operating mechanisms that turn reviews and analysis into clear decisions, owners, timelines, and sustained follow through
  • Align stakeholders and dependencies before critical launches and engineering decisions

Minimum Qualifications 

  • 5+ years of reliability engineering management experience and 7+ years of experience in software engineering, site reliability engineering, infrastructure, or platform engineering
  • Significant hands-on experience in Site Reliability Engineering, Production Engineering, or an equivalent role responsible for operating and improving production systems
  • Direct experience managing an SRE, Production Engineering, or equivalent reliability function, including ownership of its strategy, roadmap, operating model, and outcomes
  • Strong technical depth in distributed systems, cloud infrastructure, observability, and production operations
  • Experience leading high severity incident response and improving incident management practices at scale
  • Demonstrated ability to translate strategy into focused, capacity aware plans and deliver measurable outcomes
  • Track record of hiring, developing, and retaining high performing engineers and engineering leaders
  • Strong cross-functional leadership and communication, with the ability to turn complex operational data into clear decisions and drive alignment across teams

Preferred Qualifications

  • Experience operating large scale, highly available distributed systems
  • Experience implementing or evolving service-level objectives and error-budget practices
  • Experience with observability platforms such as Datadog, Grafana, Prometheus, OpenTelemetry, or similar technologies
  • Experience developing incident management, operational readiness, or resilience programs across a large engineering organization
  • Familiarity with Kubernetes, AWS, and modern cloud native architectures
  • Experience supporting major platform or architectural transitions
  • Strong product mindset when building internal reliability capabilities
  • Experience establishing executive level reliability reporting and operating reviews

Position location This role is available in the following locations: Remote

Travel requirements As a digital first company, the majority of your work can be accomplished remotely. The majority of our employees can live and work anywhere in the U.S but are encouraged to to still spend high quality time in-person collaborating via regular onsites. The in-person sessions' cadence varies depending on the team and role; most teams meet once or twice per quarter for 2-4 consecutive days at a time.

#LI-REMOTE

#LI-MidSenior


What Upstart employees say

Pay

Hours and flexibility

Workplace

Get the full story on Breakroom