1

Site Reliability Engineer Manager Jobs in Tennessee

Site Reliability Engineer

Oak Ridge, TN

$54.50 - $72.50/hr

... OLCF managed HPC clusters. Major Duties/Responsibilities: • Lead ongoing improvements in ... SRE/Systems Engineer. An equivalent combination of education and experience may be considered.

Systems Engineer - SRE Enablement

Memphis, TN · On-site

$55.50 - $73.75/hr

Participate in the incident management process, including post-mortem analysis, to continuously strengthen systemic reliability. * Run SRE training programs and reliability workshops for engineering ...

Systems Engineer - SRE Enablement

Memphis, TN

$55.25 - $73.50/hr

Participate in the incident management process, including post-mortem analysis, to continuously strengthen systemic reliability. * Run SRE training programs and reliability workshops for engineering ...

Senior Site Reliability Engineer

Knoxville, TN · On-site

$50.75 - $67.50/hr

Experience managing Linux/UNIX operating systems in a heterogeneous environment. * Solid ... to SRE/systems engineering. * An understanding of code review and familiarity with tools like ...

Site Reliability Engineer II

Nashville, TN

$55 - $73.25/hr

Kastle Systems is the leader in managed security, with a track record of introducing innovative ... Site Reliability Engineer II The SRE II sits at the intersection of software engineering and ...

Service Reliability Engineer

Nashville, TN

$55 - $73.25/hr

Incident Management & Collaboration Participate in an on-call rotation to troubleshoot and mitigate ... Partner with engineering and IT stakeholders to embed SRE best practices (SLOs, error budgets) into ...

next page

Showing results 1-20

Site Reliability Engineer Manager information

See Tennessee salary details

$9

$57

$83

How much do site reliability engineer manager jobs pay per hour?

As of Jul 28, 2026, the average hourly pay for site reliability engineer manager in Tennessee is $57.85, according to ZipRecruiter salary data. Most workers in this role earn between $49.76 and $66.11 per hour, depending on experience, location, and employer.

Will AI replace SRE jobs?

AI is expected to augment Site Reliability Engineer (SRE) roles by automating routine tasks such as monitoring, incident response, and data analysis, allowing SREs to focus on complex problem-solving and system design. While AI can improve efficiency, it is unlikely to fully replace SREs, as human expertise is essential for managing system architecture, making strategic decisions, and handling unforeseen issues.

What is a Site Reliability Engineer Manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What engineer makes $500,000 a year?

A senior or principal Site Reliability Engineer (SRE) with extensive experience, specialized skills, and often working at large tech companies or in high-cost-of-living areas can earn $500,000 or more annually. Compensation may include base salary, bonuses, and stock options, especially for those in leadership or highly technical roles. Advanced certifications and expertise in cloud platforms, automation, and system architecture are common among top earners in this field.

Is SRE a stressful job?

Site Reliability Engineer (SRE) roles can be stressful due to the high responsibility for system uptime, incident response, and maintaining service reliability. The job often involves working under pressure, handling outages, and balancing automation with manual intervention, but it also offers opportunities for skill development and process improvement. Effective SREs use monitoring tools and incident management practices to manage stress and ensure system stability.

What is the role of site reliability engineer manager?

A Site Reliability Engineer Manager oversees a team responsible for maintaining the reliability, availability, and performance of software systems. They coordinate incident response, implement automation, and ensure system scalability, often using tools like monitoring and alerting platforms. The role requires strong leadership, technical expertise, and knowledge of cloud infrastructure and DevOps practices.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How does a Site Reliability Engineer Manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What are the key skills and qualifications needed to thrive as a Site Reliability Engineer Manager, and why are they important?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.
What are the most commonly searched types of Site Reliability Engineer jobs in Tennessee? The most popular types of Site Reliability Engineer jobs in Tennessee are:
What cities in Tennessee are hiring for Site Reliability Engineer Manager jobs? Cities in Tennessee with the most Site Reliability Engineer Manager job openings:
Infographic showing various Site Reliability Engineer Manager job openings in Tennessee as of July 2026, with employment types broken down into 92% Full Time, 5% Part Time, and 3% Contract. Highlights an 89% Physical, 4% Hybrid, and 7% Remote job distribution, with an average salary of $120,335 per year, or $57.9 per hour.
Site Reliability Engineer

Site Reliability Engineer

ITR

Oak Ridge, TN

$54.50 - $72.50/hr

Full-time

Posted 6 days ago


Job description

Senior Site Reliability Engineer, HPC Infrastructure and Platforms
Overview:
Seeking highly qualified individuals to play a key role in improving the security, performance, and reliability of the HPC computing infrastructure which supports multiple highly ranked Top500 Supercomputers, including the world’s first exaflop system, Frontier.
The Team:
As a Senior Site Reliability Engineer, you will work within the HPC Infrastructure and Platforms group to support all activities of our supercomputer center. Our primary platform is the OLCF Slate Service, built on Kubernetes and Red Hat OpenShift, which provides a container orchestration service for running critical operation applications and user-managed persistent applications that run alongside our OLCF Supercomputer systems and other OLCF managed HPC clusters.
Major Duties/Responsibilities:
• Lead ongoing improvements in reliability and scalability for our Kubernetes and Linux based applications and services.
• Contribute as senior technical resource to define and implement best practices and standards for the center.
• Provide primary operational support and engineering for production applications.
• Define and implement define KPIs, processes and drive continuous improvement.
• Influence the architecture and implementation of solutions.
• Tune operating systems and applications to increase performance and reliability of services.
• Mentor junior staff and enable them for success.
• Diagnose system operational problems quickly and effectively.
• Participate in on-call rotation providing 24-hour, 7-day support and off-hours maintenance windows.
• Coordinate with vendors to resolve hardware and software problems.
• Deliver client mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote diversity, equity, inclusion, and accessibility by fostering a respectful workplace – in how we treat one another, work together, and measure success.
Basic Qualifications:
Bachelor’s Degree in computer science or closely related field and a minimum of 8 years of experience as an SRE/Systems Engineer. An equivalent combination of education and experience may be considered.
Preferred Qualifications:
• Excellent interpersonal/communication skills, and the ability to work as part of a team.
• Strong working knowledge of Unix system fundamentals and common network protocols.
• Experience managing Linux/UNIX operating systems in a heterogeneous environment.
• Solid understanding of networked computing environment concepts.
• Ability to develop and maintain programs and scripts that aid in the operation and automation using various shell (primarily bash) and high-level languages (Python or Go).
• Ability to proactively identify performance issues, problems, and areas for improvement.
• Ability to identify requirements and to define, plan, and implement requisite solutions.
• Ability to plan, organize, prioritize tasks, and complete assigned projects with minimal supervision.
• Experience with continuous integration and continuous deployment software methodologies and how they apply to SRE/systems engineering.
• An understanding of code review and familiarity with tools like GitHub and GitLab
• Experience using tools such as Nagios, Grafana and Prometheus to monitor systems, metrics, and create dashboards.
• Experience designing and implement highly available systems/services utilizing virtual machines and Kubernetes resources.
• Experience participating in an opensource community with patches accepted upstream.
• Experience deploying and maintaining automated configuration management software such as Puppet or Ansible
• Experience implementing systems-level security technologies like SELinux and following security best practices.
Special Requirement:
This position requires the ability to obtain and maintain a clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse program (WSAP) testing designed position which requires passing a pre-placement drug test and participation in an ongoing random drug testing program in which employees are subject to being randomly selected for testing. The occupant of this position will also be subject to an ongoing requirement to report any drug-related arrest or conviction or receipt of a positive drug test result.