1

Reliability Engineering Director Jobs (NOW HIRING)

$42 - $55.75/hr

Direct and manage contractor resources, including scoping work, setting quality expectations, and ... What You Will Need: * 6+ years in Site Reliability Engineering, DevOps, or Cloud Operations, with ...

Director, Site Reliability Engineering

Denver, CO · On-site

$58.75 - $78/hr

The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and observability initiatives for a portfolio of Vertafore products. This role owns SLIs/SLOs, incident response ...

Director, Site Reliability Engineering

Denver, CO · On-site

$58.75 - $78/hr

The Director, Site Reliability Engineering (SRE) will lead reliability, performance, and observability initiatives for a portfolio of Vertafore products. This role owns SLIs/SLOs, incident response ...

Manager, Site Reliability Engineering

Buffalo, NY · On-site

$55.25 - $73.50/hr

Manages an organization of employees and contingent resources through direct reports, program managers, SRE leaders, technical leads, and matrixed delivery relationships. Establishes the enterprise ...

Manager, Site Reliability Engineering

Buffalo, NY · On-site

$55.25 - $73.50/hr

Manages an organization of employees and contingent resources through direct reports, program managers, SRE leaders, technical leads, and matrixed delivery relationships. Establishes the enterprise ...

Showing results 21-40

Reliability Engineering Director information

See salary details

$73K

$194.7K

$254K

How much do reliability engineering director jobs pay per year?

As of Sep 11, 2026, the average yearly pay for reliability engineering director in the United States is $194,709.00, according to ZipRecruiter salary data. Most workers in this role earn between $141,500.00 and $253,000.00 per year, depending on experience, location, and employer.

What does a reliability engineering director do?

A Reliability Engineering Director leads teams focused on improving the reliability, availability, and performance of systems, infrastructure, or products. They set strategic goals, oversee reliability programs, and collaborate with cross-functional teams to identify risks and implement solutions. This role often involves analyzing failure data, establishing best practices, and ensuring that systems meet organizational and customer expectations for uptime and quality. Strong leadership, technical expertise, and communication skills are essential for success in this position.

What are the key skills and qualifications needed to thrive as a reliability engineering director?

To thrive as a Reliability Engineering Director, you need deep expertise in reliability principles, systems engineering, and leadership, often supported by a degree in engineering and substantial industry experience. Familiarity with tools like FMEA, Root Cause Analysis, reliability modeling software, and certifications such as CRE (Certified Reliability Engineer) are typically required. Strong communication, strategic thinking, and team management skills set top performers apart in this role. These competencies are crucial for ensuring system uptime, driving process improvements, and leading teams to deliver reliable, high-quality products or services.

What are some of the key challenges faced by a reliability engineering director in ensuring system uptime and performance?

A Reliability Engineering Director often faces the challenge of balancing rapid innovation with the need for robust, stable systems. They must anticipate and mitigate potential points of failure while collaborating with cross-functional teams to integrate reliability best practices into the development lifecycle. Additionally, they are responsible for fostering a culture of proactive incident response and continuous improvement, often leading incident post-mortems and driving long-term solutions. Managing competing priorities and aligning reliability goals with business objectives are also essential aspects of the role.

What is the difference between Reliability Engineering Director vs Reliability Engineer?

AspectReliability Engineering DirectorReliability Engineer
CredentialsBachelor's or Master's in Engineering, certifications like CRE or Six SigmaBachelor's in Engineering or related field, certifications like CRE beneficial
Work EnvironmentLeadership role overseeing teams, strategic planningTechnical role focused on analysis and testing
Employer & Industry UsageUsed in manufacturing, aerospace, energy sectors at managerial levelCommon in similar industries, focused on technical tasks

The Reliability Engineering Director focuses on strategic leadership, team management, and high-level reliability initiatives, while the Reliability Engineer handles technical analysis, testing, and implementation of reliability solutions. Both roles require relevant certifications and industry experience, but differ mainly in scope and responsibility.

What cities are hiring for Reliability Engineering Director jobs?

Cities with the most Reliability Engineering Director job openings:

What states have the most Reliability Engineering Director jobs?

States with the most job openings for Reliability Engineering Director jobs include:

What are popular job titles related to Reliability Engineering Director jobs?

For Reliability Engineering Director jobs, the most frequently searched job titles are:

Infographic showing various Reliability Engineering Director job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 90% Full Time, 6% Part Time, 2% Contract, and 1% Nights. Highlights an 83% Physical, 4% Hybrid, and 13% Remote job distribution, with an average salary of $194,709 per year, or $93.6 per hour.

Manager, Site Reliability Engineering

On-site

Delinea
IT Services • 501 - 1,000 employees

$42 - $55.75/hr

Other

Medical, Life, Retirement

Posted 5 days ago


Job description

About Delinea:

Delinea is a pioneer in securing human and machine identities through intelligent, centralized authorization, empowering organizations to seamlessly govern their interactions across the modern enterprise. Leveraging AI-powered intelligence, Delinea’s leading cloud-native Identity Security Platform applies context throughout the entire identity lifecycle - across cloud and traditional infrastructure, data, SaaS applications, and AI. It is the only platform that enables you to discover all identities including workforce, IT administrator, developers, and machines assign appropriate access levels, detect irregularities, and respond to threats in real-time. With deployment in weeks, not months, 90% fewer resources to manage than the nearest competitor, and a 99.995% uptime, Delinea delivers robust security and operational efficiency without compromise. Learn more about Delinea on Delinea.com, LinkedIn, X, and YouTube.

Summary:

Delinea is looking for a hands-on Manager of Site Reliability Engineering to lead the SRE and DevOps engineers supporting the Delinea products. This is a working manager role. You will be expected to lead people and lead work: writing and reviewing automation, digging into AKS and Azure telemetry, commanding Sev1 and Sev2 incidents, and improving the observability and deployment practices your team depends on.

The initial scope is the Platform SRE and DevOps team. Over time, the role is expected to expand to cover our FedRAMP High environment and the broader Commercial Platform footprint, so comfort operating in a regulated environment and a willingness to grow scope are essential.

You will lead a blended team of full-time engineers and contractors distributed across multiple time zones. Meeting your team inside their working hours is an expectation of the role. On-call participation is required.

What You Will Do:
  • Lead hands-on. Spend a meaningful portion of your week in the environment: reviewing pull requests, validating pipeline changes, tuning monitors and dashboards, running queries in Datadog, and troubleshooting production issues alongside your engineers. This role does not sit above the work.

  • Own availability and performance of the Delinea Platform production environments across Azure and AWS, including AKS workloads, ingress and networking, data services, messaging, and CDN or WAF layers.

  • Manage a blended team. Hire, onboard, coach, and develop full-time SRE engineers. Direct and manage contractor resources, including scoping work, setting quality expectations, and reviewing deliverables.

  • Lead a distributed team. Run one-on-ones, standups, and planning sessions at times that work for engineers in other geographies.

  • Participate in on-call. Carry the pager as part of the rotation, act as incident commander for Sev1 and Sev2 events, drive engagement of the right responders, and own communication cadence with support, engineering, and leadership until resolution.

  • Command incident response end to end. Own detection, triage, mitigation, customer-facing status communication, and post-incident review. Ensure RCAs are written to a customer-ready standard, preventative actions have owners and target dates, and those actions are driven to closure.

  • Raise the observability bar. Improve detection coverage so that issues are found by our monitoring rather than by a customer ticket. Own SLI and SLO definition, alert quality and noise reduction, synthetic coverage, APM instrumentation, log hygiene, and dashboard standards.

  • Support FedRAMP and regulated operations. Grow into supporting our FedRAMP High environment, including change control discipline, evidence collection, boundary awareness, and the operational differences between government and commercial environments.

  • Reduce toil through automation. Set the expectation that repeat manual work becomes code. Prioritize automation backlog alongside project and reliability work.

  • Report on operational health. Produce and present incident metrics, trends, and reliability commitments to leadership, and translate them into a concrete improvement plan.

What You Will Need:
  • 6+ years in Site Reliability Engineering, DevOps, or Cloud Operations, with demonstrated ownership of production SaaS systems.
  • 2+ years of direct people leadership, including performance management, hiring, and coaching. Experience managing contractors or an outsourced delivery team is desired.
  • Current, hands-on production experience with the Delinea technology stack, including Azure Kubernetes Service, core Azure services (SQL, Redis, Service Bus, Blob Storage), AWS services (SES, EC2, RDS), WAF, Azure DevOps pipelines, Datadog, and Atlassian Jira Service Management.
  • Hands-on experience across both Azure and AWS is required. You should be able to administer, troubleshoot, and reason about cost and security posture in each.
  • Deep observability expertise. Demonstrated ownership of an observability framework at scale: metrics, logs, traces, synthetics, SLOs, and alerting strategy. Hands-on proficiency with Datadog or an equivalent platform, including APM trace analysis and log-based troubleshooting.
  • Proven incident command. You have run major incidents as the incident commander, coordinated multiple responders under pressure, communicated to customers and executives during active impact, and authored the RCA afterward.
  • Strong cloud networking and security fundamentals: load balancing, DNS, TLS and certificate lifecycle, firewalls, VPN, routing, and identity and access management.
  • Automation and scripting ability in PowerShell, Python, Bash, or similar, plus practical infrastructure-as-code experience (Terraform, ARM, or Bicep).
  • Practical experience with multi-region, multi-tenant SaaS architectures, including backup, redundancy, and disaster recovery approaches.
  • Excellent written communication. You will write and approve customer-facing status updates and incident summaries under time pressure.
  • Willingness and availability to work across time zones and to participate in an on-call rotation.
We Would Love to See
  • Direct experience operating in a FedRAMP or other regulated environment (Azure Government, IL4/IL5, SOC 2, ISO 27001).
  • Experience standing up or maturing an incident management program, including sev definitions, escalation paths, on-call structure, and post-incident review process.
  • Experience with public status page operations and customer notification practices.
  • Experience with Atlassian Jira Service Management, Confluence, and Azure DevOps as the operational toolchain.
  • Track record of reducing customer-detected incidents through improved monitoring coverage.
  • Cost optimization experience across Azure and AWS on a meaningful scale.
Why work at Delinea?
  • We're passionate problem-solvers helping the world's largest organizations protect what matters most: their human and machine identities.
  • We invest in people who are smart, self-motivated, and collaborative.
  • What we offer in return is meaningful work, a culture of innovation and great career progression.
At Delinea, our core values are STRONG and guide our behaviors and success:
  • Spirited - We bring energy and passion to everything we do
  • Trust - We act with integrity and deliver on our commitments
  • Respect - We listen, value different perspectives, and work as one team
  • Ownership - We take initiative and follow through
  • Nimble - We adapt quickly in a fast-changing environment
  • Global - We embrace diverse people and ideas to drive better outcomes

We believe weaving these core values into our day-to-day actions, and our process for hiring, evaluating, and promoting employees, helps us cultivate a work environment that embraces collaboration and camaraderie.

We take care of our employees. We offer competitive salaries, a meaningful bonus program, and excellent benefits, including healthcare insurance, as well as pension/retirement matching, comprehensive life insurance, an employee assistance program, time off plans, and paid company holidays.

Delinea is an Equal Opportunity and Affiliated employer and prohibits discrimination and harassment of any type with regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

Upon conditional offer of employment, candidates are required to complete comprehensive criminal background check, verification of education, and verification of employment, per employment policy. In addition, all publicly posted social media sites may be reviewed.

#J-18808-Ljbffr