1

Site Reliability Engineer Manager Jobs in Oak Ridge, TN

Site Reliability Engineer

Oak Ridge, TN · On-site

$54.50 - $72.50/hr

Experience managing Linux/UNIX operating systems in a heterogeneous environment. * Solid ... to SRE/systems engineering. * An understanding of code review and familiarity with tools like ...

Reliability Engineer

Oak Ridge, TN · On-site

$98K - $123K/yr

Fill a critical role in optimizing asset management processes through utilization of proactive ... Site The NNSA's Y-12 National Security Complex, in Oak Ridge, Tennessee, is the nation's only ...

Reliability Engineer

Alcoa, TN · On-site

$90 - $120/hr

Actively attend departmental daily management and maintenance meetings as required. * Identify and share best practices in equipment reliability and maintenance. * Train engineers and maintenance ...

New

Reliability Engineer, Level II

Oak Ridge, TN · On-site

$98K - $123K/yr

Pay, benefits, Human Resources, and Employment Management functions are provided by AIS. Qualified ... Interact with organizations across the Site. * Present information on clearly and effectively ...

Work with scientific and technical users to help them use Kubernetes Basic Qualifications: * 5+ years of experience working as an SRE/Systems Administrator/Systems Engineer * Bachelor's degree or an ...

ENGINEER AUTOMATION RELIABILITY

Knoxville, TN · On-site

$85K - $108K/yr

The Reliability & Projects Engineer will be responsible for the identification, development, and ... Proven experience as a self-starter who effectively manages multiple tasks and achieves results ...

ENGINEER AUTOMATION RELIABILITY

Knoxville, TN · On-site

$85K - $108K/yr

The Reliability & Projects Engineer will be responsible for the identification, development, and ... Proven experience as a self-starter who effectively manages multiple tasks and achieves results ...

Work in close partnership with the site Reliability Engineer to analyze equipment history and ... Exceptional organizational and time-management skills with the ability to manage a large backlog ...

Manufacturing Engineer Manager

Clinton, TN · On-site

$90K - $111K/yr

... Reliability & Maintenance Support - Improve equipment uptime and manufacturing efficiency through engineering and maintenance collaboration. - Lead downtime reduction activities, root cause analysis ...

Work in close partnership with the site Reliability Engineer to analyze equipment history and ... Exceptional organizational and time-management skills with the ability to manage a large backlog ...

Work in close partnership with the site Reliability Engineer to analyze equipment history and ... Exceptional organizational and time-management skills with the ability to manage a large backlog ...

next page

Showing results 1-20

Site Reliability Engineer Manager information

See Oak Ridge, TN salary details

$10

$60

$87

How much do site reliability engineer manager jobs pay per hour?

As of Aug 8, 2026, the average hourly pay for site reliability engineer manager in Oak Ridge, TN is $60.94, according to ZipRecruiter salary data. Most workers in this role earn between $52.40 and $69.62 per hour, depending on experience, location, and employer.

What is a site reliability engineer manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How does a site reliability engineer manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What are the key skills and qualifications needed to thrive as a site reliability engineer manager?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.
What are the most commonly searched types of Site Reliability Engineer jobs in Oak Ridge, TN? The most popular types of Site Reliability Engineer jobs in Oak Ridge, TN are:
What cities near Oak Ridge, TN are hiring for Site Reliability Engineer Manager jobs? Cities near Oak Ridge, TN with the most Site Reliability Engineer Manager job openings:

Site Reliability Engineer

ITR

Oak Ridge, TN • On-site

$54.50 - $72.50/hr

Full-time

Posted 17 days ago


Job description

Senior Site Reliability Engineer, HPC Infrastructure and Platforms
Overview:
Seeking highly qualified individuals to play a key role in improving the security, performance, and reliability of the HPC computing infrastructure which supports multiple highly ranked Top500 Supercomputers, including the world’s first exaflop system, Frontier.
The Team:
As a Senior Site Reliability Engineer, you will work within the HPC Infrastructure and Platforms group to support all activities of our supercomputer center. Our primary platform is the OLCF Slate Service, built on Kubernetes and Red Hat OpenShift, which provides a container orchestration service for running critical operation applications and user-managed persistent applications that run alongside our OLCF Supercomputer systems and other OLCF managed HPC clusters.
Major Duties/Responsibilities:
• Lead ongoing improvements in reliability and scalability for our Kubernetes and Linux based applications and services.
• Contribute as senior technical resource to define and implement best practices and standards for the center.
• Provide primary operational support and engineering for production applications.
• Define and implement define KPIs, processes and drive continuous improvement.
• Influence the architecture and implementation of solutions.
• Tune operating systems and applications to increase performance and reliability of services.
• Mentor junior staff and enable them for success.
• Diagnose system operational problems quickly and effectively.
• Participate in on-call rotation providing 24-hour, 7-day support and off-hours maintenance windows.
• Coordinate with vendors to resolve hardware and software problems.
• Deliver client mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote diversity, equity, inclusion, and accessibility by fostering a respectful workplace – in how we treat one another, work together, and measure success.
Basic Qualifications:
Bachelor’s Degree in computer science or closely related field and a minimum of 8 years of experience as an SRE/Systems Engineer. An equivalent combination of education and experience may be considered.
Preferred Qualifications:
• Excellent interpersonal/communication skills, and the ability to work as part of a team.
• Strong working knowledge of Unix system fundamentals and common network protocols.
• Experience managing Linux/UNIX operating systems in a heterogeneous environment.
• Solid understanding of networked computing environment concepts.
• Ability to develop and maintain programs and scripts that aid in the operation and automation using various shell (primarily bash) and high-level languages (Python or Go).
• Ability to proactively identify performance issues, problems, and areas for improvement.
• Ability to identify requirements and to define, plan, and implement requisite solutions.
• Ability to plan, organize, prioritize tasks, and complete assigned projects with minimal supervision.
• Experience with continuous integration and continuous deployment software methodologies and how they apply to SRE/systems engineering.
• An understanding of code review and familiarity with tools like GitHub and GitLab
• Experience using tools such as Nagios, Grafana and Prometheus to monitor systems, metrics, and create dashboards.
• Experience designing and implement highly available systems/services utilizing virtual machines and Kubernetes resources.
• Experience participating in an opensource community with patches accepted upstream.
• Experience deploying and maintaining automated configuration management software such as Puppet or Ansible
• Experience implementing systems-level security technologies like SELinux and following security best practices.
Special Requirement:
This position requires the ability to obtain and maintain a clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse program (WSAP) testing designed position which requires passing a pre-placement drug test and participation in an ongoing random drug testing program in which employees are subject to being randomly selected for testing. The occupant of this position will also be subject to an ongoing requirement to report any drug-related arrest or conviction or receipt of a positive drug test result.