1

Sre Director Jobs (NOW HIRING)

Showing results 41-60

Sre Director information

What is an SRE director?

SRE Directors are senior leaders responsible for overseeing Site Reliability Engineering (SRE) teams within an organization. They set strategic direction, manage large-scale reliability initiatives, and ensure that systems are highly available, scalable, and efficient. SRE Directors collaborate closely with engineering, product, and operations teams to define service level objectives, incident management processes, and reliability best practices. Their role is pivotal in balancing the needs of innovation, reliability, and operational excellence across the company.

How does an SRE director typically collaborate with other engineering and business leaders within an organization?

As an SRE Director, you will regularly interact with engineering, product, and business leadership to ensure system reliability aligns with organizational goals. Collaboration often involves setting reliability standards, negotiating service-level objectives, and prioritizing initiatives that balance feature delivery with operational stability. You'll also oversee cross-functional incident response and postmortems, championing a culture of continuous improvement. Effective communication and stakeholder management are essential to align technical efforts with business needs.

What are the key skills and qualifications needed to thrive as an SRE director, and why are they important?

To thrive as an SRE Director, you need deep expertise in site reliability engineering, infrastructure management, and a solid background in computer science or a related field, often supported by relevant leadership experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), automation tools, monitoring systems, and certifications such as Google Professional SRE or AWS Certified DevOps Engineer is common. Outstanding communication, strategic thinking, and team leadership are crucial soft skills that drive organizational reliability and performance. These abilities are essential to ensure system uptime, scalable operations, and to foster a culture of reliability and continuous improvement across engineering teams.

Is SRE a good career path?

Site Reliability Engineering (SRE) is a growing field that combines software engineering and systems administration to ensure the reliability and scalability of services. It typically requires skills in coding, automation, and monitoring tools, and offers strong job growth and competitive salaries. Many organizations value SRE roles for their focus on system stability and incident response.
More about Sre Director jobs

What cities are hiring for Sre Director jobs?

Cities with the most Sre Director job openings:

What are the most commonly searched types of Sre jobs?

The most popular types of Sre jobs are:

What states have the most Sre Director jobs?

States with the most job openings for Sre Director jobs include:

Infographic showing various Sre Director job openings in the United States as of August 2026, with employment types broken down into 2% As Needed, 83% Full Time, 12% Part Time, 1% Temporary, and 2% Contract. Highlights an 91% Physical, 3% Hybrid, and 6% Remote job distribution.

$58 - $77.25/hr

Full-time

Posted 7 days ago


Job description

Overview

We are seeking an experienced Site Reliability Engineer (SRE) – Microsoft Hyper-V & Private Cloud to operate highly available private cloud and Virtual Desktop Infrastructure (VDI) platforms based on Microsoft Hyper-V. This role combines deep Hyper-V expertise with modern Site Reliability Engineering practices to improve platform reliability, scalability, performance, automation, and operational excellence.

The ideal candidate will be responsible for ensuring platform reliability through infrastructure automation, OS upgrades, proactive monitoring, incident response, capacity planning, and continuous service improvement while collaborating with infrastructure, security, and application teams.

Key Responsibilities

  • Operate enterprise-scale private cloud infrastructure built on Microsoft Hyper-V.
  • Optimize, and support highly available VDI environments on Hyper-V.
  • Improve platform reliability, availability, scalability, and resiliency by applying SRE principles and engineering best practices.
  • Disaster recovery, backup, patch management, and business continuity strategies.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational metrics for critical infrastructure services.
  • Automate infrastructure provisioning, configuration management, and operational workflows using PowerShell and Infrastructure as Code (IaC) principles wherever applicable.
  • Manage Hyper-V Failover Clusters, host lifecycle, storage, networking, and capacity to ensure high availability and business continuity.
  • Develop proactive monitoring, alerting, logging, and observability capabilities to detect and prevent service degradation.
  • Lead incident response for infrastructure-related outages, perform root cause analysis (RCA), and implement preventive actions through post-incident reviews.
  • Perform capacity planning, performance tuning, and resource optimization across Hyper-V clusters and VDI platforms.
  • Support infrastructure migration initiatives including P2V, V2V, workload modernization, and private cloud transformations.
  • Collaborate closely with Security, Networking, Platform Engineering, and Application teams to improve platform reliability and operational efficiency.
  • Develop and maintain technical documentation, architecture diagrams, operational runbooks, automation scripts, and standard operating procedures.
  • Mentor junior engineers and promote SRE culture, automation, and operational best practices across the team.


Experience & Qualifications

  • 6+ years of infrastructure engineering experience with at least 4+ years of hands-on Microsoft Hyper-V administration.
  • Demonstrated experience operating mission-critical enterprise infrastructure with high availability and reliability requirements.
  • Proven experience implementing automation to reduce operational overhead and improve service reliability.
  • Experience supporting enterprise private cloud and VDI environments.
  • Experience participating in incident response, root cause analysis, and continuous service improvement initiatives.
  • Microsoft certifications such as Microsoft Certified: Windows Server Hybrid Administrator Associate or equivalent are desirable.
  • Experience in Banking or Financial Services environments is advantageous.

    Preferred Skills

    • Windows Server 2016/2019/2022 administration.
    • Experience with backup and disaster recovery solutions such as Veeam, Altaro, or native Hyper-V Replica.
    • Exposure to hybrid cloud and private cloud platforms.
    • Familiarity with monitoring and observability platforms such as SCOM, Azure Monitor, Prometheus, Grafana, Splunk, or similar tools.
    • Experience supporting enterprise VDI environments.
    • Understanding of ITIL Incident, Problem, Change, and Release Management.
    • Experience working in regulated industries such as Banking or Financial Services.