1

Site Reliability Engineer Manager Jobs in Tennessee

Systems Engineer - SRE Enablement

Memphis, TN · On-site

$55.50 - $73.75/hr

Participate in the incident management process, including post-mortem analysis, to continuously strengthen systemic reliability. * Run SRE training programs and reliability workshops for engineering ...

Systems Engineer - SRE Enablement

Memphis, TN · On-site

$55.25 - $73.50/hr

Participate in the incident management process, including post-mortem analysis, to continuously strengthen systemic reliability. * Run SRE training programs and reliability workshops for engineering ...

Site Reliability Engineer II

Nashville, TN

$55 - $73.25/hr

Kastle Systems is the leader in managed security, with a track record of introducing innovative ... Site Reliability Engineer II The SRE II sits at the intersection of software engineering and ...

Senior Site Reliability Engineer

Knoxville, TN

$50.75 - $67.50/hr

Experience managing Linux/UNIX operating systems in a heterogeneous environment. * Solid ... to SRE/systems engineering. * An understanding of code review and familiarity with tools like ...

New

Principal Site Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

Incident and Service Lifecycle Management: - Exercises judgment when performing data collection ... site reliability trends, sharing valuable insights and information with senior team members ...

Principal Site Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

Incident and ServiceLifecycle Management: - Exercisesjudgment when performing data collection ... site reliability trends, sharing valuableinsights and information with senior team members ...

Incident and ServiceLifecycle Management: - Exercisesjudgment when performing data collection ... site reliability trends, sharing valuableinsights and information with senior team members ...

Service Reliability Engineer

Nashville, TN · On-site

$55 - $73.25/hr

Incident Management & Collaboration Participate in an on-call rotation to troubleshoot and mitigate ... Partner with engineering and IT stakeholders to embed SRE best practices (SLOs, error budgets) into ...

next page

Showing results 1-20

Site Reliability Engineer Manager information

See Tennessee salary details

$9

$57

$83

How much do site reliability engineer manager jobs pay per hour?

As of Jul 20, 2026, the average hourly pay for site reliability engineer manager in Tennessee is $57.85, according to ZipRecruiter salary data. Most workers in this role earn between $49.76 and $66.11 per hour, depending on experience, location, and employer.

Will AI replace SRE jobs?

AI is expected to augment Site Reliability Engineer (SRE) roles by automating routine tasks such as monitoring, incident response, and data analysis, allowing SREs to focus on complex problem-solving and system design. While AI can improve efficiency, it is unlikely to fully replace SREs, as human expertise is essential for managing system architecture, making strategic decisions, and handling unforeseen issues.

What is a Site Reliability Engineer Manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What engineer makes $500,000 a year?

A senior or principal Site Reliability Engineer (SRE) with extensive experience, specialized skills, and often working at large tech companies or in high-cost-of-living areas can earn $500,000 or more annually. Compensation may include base salary, bonuses, and stock options, especially for those in leadership or highly technical roles. Advanced certifications and expertise in cloud platforms, automation, and system architecture are common among top earners in this field.

Is SRE a stressful job?

Site Reliability Engineer (SRE) roles can be stressful due to the high responsibility for system uptime, incident response, and maintaining service reliability. The job often involves working under pressure, handling outages, and balancing automation with manual intervention, but it also offers opportunities for skill development and process improvement. Effective SREs use monitoring tools and incident management practices to manage stress and ensure system stability.

What is the role of site reliability engineer manager?

A Site Reliability Engineer Manager oversees a team responsible for maintaining the reliability, availability, and performance of software systems. They coordinate incident response, implement automation, and ensure system scalability, often using tools like monitoring and alerting platforms. The role requires strong leadership, technical expertise, and knowledge of cloud infrastructure and DevOps practices.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How does a Site Reliability Engineer Manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What are the key skills and qualifications needed to thrive as a Site Reliability Engineer Manager, and why are they important?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.
What are the most commonly searched types of Site Reliability Engineer jobs in Tennessee? The most popular types of Site Reliability Engineer jobs in Tennessee are:
What cities in Tennessee are hiring for Site Reliability Engineer Manager jobs? Cities in Tennessee with the most Site Reliability Engineer Manager job openings:
Infographic showing various Site Reliability Engineer Manager job openings in Tennessee as of July 2026, with employment types broken down into 93% Full Time, 2% Part Time, 1% Temporary, and 4% Contract. Highlights an 87% Physical, 4% Hybrid, and 9% Remote job distribution, with an average salary of $120,335 per year, or $57.9 per hour.

Site Reliability Engineer

Everest Global Solutions

Brentwood, TN • On-site

$54 - $71.75/hr

Other

Posted 11 days ago


Job description

Required Technical

Expertise:

• 10+ years of hands-on experience in Java application development,

production engineering, and Site Reliability Engineering (SRE).

• Strong expertise in Core Java, Java 11/17+, Spring Boot, Spring Cloud,

Microservices, JVM Internals, Garbage Collection (GC) tuning, and

multithreading.

• Extensive experience with Google Cloud Platform (Google Cloud Platform), including Google

Kubernetes Engine (GKE), Compute Engine, Cloud Storage, Cloud SQL,
Pub/Sub, IAM, VPC, Cloud Monitoring, Cloud Logging, Secret Manager, and

Cloud Load Balancing.

• Expert-level experience administering and troubleshooting Kubernetes

production environments, including cluster management, networking,

autoscaling, RBAC, storage, Helm, Ingress Controllers, and service mesh

technologies.

• Hands-on experience with Docker, container orchestration, and

Infrastructure as Code (IaC) using Terraform.

• Strong experience building and maintaining enterprise CI/CD pipelines

using GitLab CI/CD, Jenkins, GitOps, and cloud-native deployment tools.

• Experience implementing and supporting zero-downtime deployment

strategies, including Blue-Green and Canary deployments.

• Strong proficiency in Linux/Unix administration, shell scripting (Bash), and

operational automation using Python or Go.

• Hands-on experience with Kafka, Kafka Streams, Pub/Sub, or other event-

driven messaging platforms supporting high-volume distributed systems.

• Experience implementing enterprise monitoring, logging, and observability

solutions using Prometheus, Grafana, Datadog, Splunk, OpenTelemetry,

Cloud Monitoring, and Kiali.

• Strong understanding of networking concepts, including TCP/IP, DNS, Load

Balancers, NGINX, Ingress Controllers, TLS/SSL, and Service Mesh (Istio).

• Experience supporting highly available, scalable, and mission-critical

production environments with 24x7 operational responsibilities.

• Strong experience performing production incident management, RCA

(Root Cause Analysis), performance tuning, capacity planning, and

reliability engineering.

• Experience with enterprise security best practices, including IAM, Secret

Manager, RBAC, least-privilege access, vulnerability remediation, and

container security.

• Experience supporting environments compliant with PCI-DSS, SOC2, SOX,

ISO 27001, or similar regulatory frameworks.

• Prior experience supporting large-scale retail, eCommerce, omnichannel,

supply chain, order management, inventory management, or payment

processing platforms.

• Excellent troubleshooting, analytical, communication, and stakeholder

management skills.

• Experience with Anthos, ArgoCD, Cloud Build, Cloud Deploy, Redis,

PostgreSQL, MongoDB, and Cassandra.

• Experience with eBPF, distributed tracing, performance engineering, and

advanced observability.

• Exposure to VMware, disaster recovery planning, chaos engineering, and

multi-region Kubernetes deployments.

• Experience leading SRE initiatives, mentoring engineering teams, and

driving operational excellence in enterprise environments.

Certifications

Required:

• Google Professional Cloud Architect OR Google Professional Cloud DevOps

Engineer

• Certified Kubernetes Administrator (CKA) OR Certified Kubernetes Security