1

Sre Manager Jobs in Tennessee (NOW HIRING)

Site Reliability Engineer

Oak Ridge, TN

$54.50 - $72.50/hr

Experience managing Linux/UNIX operating systems in a heterogeneous environment. * Solid ... to SRE/systems engineering. * An understanding of code review and familiarity with tools like ...

Systems Engineer - SRE Enablement

Memphis, TN · On-site

$55.50 - $73.75/hr

Participate in the incident management process, including post-mortem analysis, to continuously strengthen systemic reliability. * Run SRE training programs and reliability workshops for engineering ...

Systems Engineer - SRE Enablement

Memphis, TN · On-site

$55.25 - $73.50/hr

Participate in the incident management process, including post-mortem analysis, to continuously strengthen systemic reliability. * Run SRE training programs and reliability workshops for engineering ...

Senior Site Reliability Engineer

Knoxville, TN · On-site

$50.75 - $67.50/hr

Experience managing Linux/UNIX operating systems in a heterogeneous environment. * Solid ... to SRE/systems engineering. * An understanding of code review and familiarity with tools like ...

Site Reliability Engineer II

Nashville, TN · On-site

$55 - $73.25/hr

Kastle Systems is the leader in managed security, with a track record of introducing innovative ... Site Reliability Engineer II The SRE II sits at the intersection of software engineering and ...

Senior Site Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

Experience with infrastructure-as-code or configuration-management tools. * Production support, incident response, or SRE operational experience. * Knowledge of monitoring, alerting, centralized ...

You've managed ElasticSearch at scale and know the difference between logs that help and logs that just cost money. You think in terms of infrastructure as code and believe Terraform plans should be ...

You've managed ElasticSearch at scale and know the difference between logs that help and logs that just cost money. You think in terms of infrastructure as code and believe Terraform plans should be ...

You've managed ElasticSearch at scale and know the difference between logs that help and logs that just cost money. You think in terms of infrastructure as code and believe Terraform plans should be ...

Site Reliability Engineer 2

Nashville, TN · On-site

$55 - $73.25/hr

Three or more years of experience in site reliability engineering, systems administration ... Experience with ticketing, incident-management, or change-management systems. * Strong ...

Site Reliability Engineer 2

Nashville, TN · On-site

$55 - $73.25/hr

Three or more years of experience in site reliability engineering, systems administration ... Experience with ticketing, incident-management, or change-management systems. * Strong ...

Senior Site Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

Three or more years of experience in site reliability engineering, systems administration ... Experience with ticketing, incident-management, or change-management systems. * Strong ...

Principal Site Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud ... Experience managing complex or high-risk production changes. Troubleshooting and Operational ...

Lead Principal Site Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

Define and drive the site reliability engineering strategy for large-scale, distributed, and ... Incident Management and Problem Resolution * Provide technical leadership during complex, high ...

Principal Site Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

Automate knowledge management and operational documentation. * Enhance developer productivity and ... Site Reliability Engineering. * Evaluate emerging technologies and drive adoption where they ...

Principal Site Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

Extensive experience in site reliability engineering, systems engineering, infrastructure ... Experience managing complex or high-risk production changes. Troubleshooting and Operational ...

next page

Showing results 1-20

Sre Manager information

See Tennessee salary details

$56.3K

$106.6K

$152.9K

How much do sre manager jobs pay per year?

As of Aug 19, 2026, the average yearly pay for sre manager in Tennessee is $106,634.00, according to ZipRecruiter salary data. Most workers in this role earn between $85,800.00 and $127,100.00 per year, depending on experience, location, and employer.

What is an SRE manager?

SRE Managers are leaders responsible for overseeing Site Reliability Engineering (SRE) teams. They ensure the reliability, scalability, and performance of software systems by guiding engineers in implementing best practices, automation, and monitoring processes. SRE Managers collaborate closely with development and operations teams to balance feature development with system stability. Their role also includes mentoring SREs, managing incident response, and driving improvements in system reliability and operational efficiency.

What are some common challenges SRE managers face when leading Site Reliability Engineering teams?

SRE Managers often encounter challenges balancing reliability with rapid development, ensuring their teams have the right mix of software engineering and operations skills. They must also foster a culture of continuous improvement while managing on-call rotations and incident response without causing burnout. Additionally, collaborating effectively with development and product teams to set realistic service level objectives (SLOs) and drive adoption of SRE best practices can require strong communication and negotiation skills.

What are the key skills and qualifications needed to thrive as an SRE manager, and why are they important?

To thrive as an SRE Manager, you need a deep understanding of site reliability engineering principles, strong experience with systems architecture, and a background in computer science or a related field. Familiarity with tools such as Kubernetes, Prometheus, cloud platforms, and CI/CD pipelines, as well as certifications like AWS Certified Solutions Architect, are commonly expected. Leadership, effective communication, and problem-solving skills are crucial for driving team performance and collaborating across departments. These skills ensure high system reliability, efficient incident management, and a culture of continuous improvement within technical organizations.

What is the difference between Sre Manager vs DevOps Engineer?

AspectSre ManagerDevOps Engineer
CredentialsTypically requires a Bachelor's/Master's in CS or related field, with certifications like AWS, Google Cloud, or KubernetesSimilar credentials, often with cloud certifications and scripting skills
Work EnvironmentLeads teams, manages incident response, and oversees reliability strategiesFocuses on automation, CI/CD pipelines, and infrastructure deployment
Industry UsageCommon in large tech companies, financial services, and cloud providersWidely used across startups, tech firms, and enterprises adopting DevOps practices

The Sre Manager and DevOps Engineer roles share overlapping skills in cloud computing, automation, and infrastructure management. While the Sre Manager oversees reliability and team coordination, the DevOps Engineer focuses on implementing automation tools and deployment pipelines. Both roles are crucial for modern IT operations, but the Sre Manager typically has a broader leadership responsibility, whereas the DevOps Engineer is more hands-on with technical implementation.

What does an SRE manager do?

An SRE manager oversees the reliability and performance of software systems, leading teams that implement automation, monitoring, and incident response processes. They coordinate efforts to ensure system availability, scalability, and efficiency, often using tools like monitoring dashboards and incident management platforms. Strong leadership, technical expertise, and understanding of service level objectives are essential in this role.

What are the most commonly searched types of Sre jobs in Tennessee?

The most popular types of Sre jobs in Tennessee are:

What job categories do people searching Sre Manager jobs in Tennessee look for?

The top searched job categories for Sre Manager jobs in Tennessee are:

What cities in Tennessee are hiring for Sre Manager jobs?

Cities in Tennessee with the most Sre Manager job openings:

Infographic showing various Sre Manager job openings in Tennessee as of August 2026, with employment types broken down into 46% Full Time, and 54% Contract. Highlights an 74% In-person, and 26% Remote job distribution, with an average salary of $106,634 per year, or $51.3 per hour.

Site Reliability Engineer

ITR

Oak Ridge, TN

$54.50 - $72.50/hr

Full-time

Posted 28 days ago


Job description

Senior Site Reliability Engineer, HPC Infrastructure and Platforms
Overview:
Seeking highly qualified individuals to play a key role in improving the security, performance, and reliability of the HPC computing infrastructure which supports multiple highly ranked Top500 Supercomputers, including the world’s first exaflop system, Frontier.
The Team:
As a Senior Site Reliability Engineer, you will work within the HPC Infrastructure and Platforms group to support all activities of our supercomputer center. Our primary platform is the OLCF Slate Service, built on Kubernetes and Red Hat OpenShift, which provides a container orchestration service for running critical operation applications and user-managed persistent applications that run alongside our OLCF Supercomputer systems and other OLCF managed HPC clusters.
Major Duties/Responsibilities:
• Lead ongoing improvements in reliability and scalability for our Kubernetes and Linux based applications and services.
• Contribute as senior technical resource to define and implement best practices and standards for the center.
• Provide primary operational support and engineering for production applications.
• Define and implement define KPIs, processes and drive continuous improvement.
• Influence the architecture and implementation of solutions.
• Tune operating systems and applications to increase performance and reliability of services.
• Mentor junior staff and enable them for success.
• Diagnose system operational problems quickly and effectively.
• Participate in on-call rotation providing 24-hour, 7-day support and off-hours maintenance windows.
• Coordinate with vendors to resolve hardware and software problems.
• Deliver client mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote diversity, equity, inclusion, and accessibility by fostering a respectful workplace – in how we treat one another, work together, and measure success.
Basic Qualifications:
Bachelor’s Degree in computer science or closely related field and a minimum of 8 years of experience as an SRE/Systems Engineer. An equivalent combination of education and experience may be considered.
Preferred Qualifications:
• Excellent interpersonal/communication skills, and the ability to work as part of a team.
• Strong working knowledge of Unix system fundamentals and common network protocols.
• Experience managing Linux/UNIX operating systems in a heterogeneous environment.
• Solid understanding of networked computing environment concepts.
• Ability to develop and maintain programs and scripts that aid in the operation and automation using various shell (primarily bash) and high-level languages (Python or Go).
• Ability to proactively identify performance issues, problems, and areas for improvement.
• Ability to identify requirements and to define, plan, and implement requisite solutions.
• Ability to plan, organize, prioritize tasks, and complete assigned projects with minimal supervision.
• Experience with continuous integration and continuous deployment software methodologies and how they apply to SRE/systems engineering.
• An understanding of code review and familiarity with tools like GitHub and GitLab
• Experience using tools such as Nagios, Grafana and Prometheus to monitor systems, metrics, and create dashboards.
• Experience designing and implement highly available systems/services utilizing virtual machines and Kubernetes resources.
• Experience participating in an opensource community with patches accepted upstream.
• Experience deploying and maintaining automated configuration management software such as Puppet or Ansible
• Experience implementing systems-level security technologies like SELinux and following security best practices.
Special Requirement:
This position requires the ability to obtain and maintain a clearance from the Department of Energy. As such, this position is a Workplace Substance Abuse program (WSAP) testing designed position which requires passing a pre-placement drug test and participation in an ongoing random drug testing program in which employees are subject to being randomly selected for testing. The occupant of this position will also be subject to an ongoing requirement to report any drug-related arrest or conviction or receipt of a positive drug test result.