1

Site Reliability Engineer Manager Jobs in Florida

Site Reliability Engineering Lead

Boca Raton, FL · On-site

$54 - $72/hr

Experience leading SRE Team/s, including rotating staff across products / platforms, mentoring and development, objective setting and general line management * Experience building and operating ...

Staff Site Reliability Engineer

FL · On-site +1

$55.25 - $73.25/hr

Strong understanding of modern SRE principles, including reliability engineering, observability, toil reduction, incident management, and error budgets. * Strong scripting or programming experience ...

Site Reliability Engineer

Tampa, FL · On-site

$53.75 - $71.50/hr

\n \n \n Looking for an SRE experienced in monitoring and observability using Prometheus, Grafana, ELK Stack, DataDog, and Sentry. You will assist in converting clients off of Datadog to an Otel ...

$105 - $175/hr

Lead incident management processes, including post-mortems and Root Cause Analysis (RCA). * Drive ... Contribute to SRE capability building through training and coaching. Requirements * Experience ...

The SRE will partner with database administration and SRE teams to ensure the reliability, security ... Lead incident management processes, including post-mortems and Root Cause Analysis (RCA). * Drive ...

SRE AIOps Engineer

Miami, FL · On-site

$54.50 - $72.50/hr

Implement and manage Infrastructure as Code (Terraform, CloudFormation) for SRE tooling and observability. * Work with internal AI tools and Atlassian Rovo to advance AI-driven operational ...

New

Sr. Site Reliability Engineer

Orlando, FL · On-site

$53.25 - $70.75/hr

As a Senior Site Reliability Engineer, you will serve as a technical leader within the platform engineering ecosystem, helping shape engineering standards, drive reliability initiatives, and ...

SRE AI Ops Engineer

Miami, FL · On-site

$54.50 - $72.50/hr

Create and manage Infrastructure as Code solutions for observability and SRE tooling * Monitor production environments and proactively identify operational issues * Support AEM and other enterprise ...

Site Reliability Engineer II

Orlando, FL · On-site

$53.25 - $70.75/hr

Kastle Systems is the leader in managed security, with a track record of introducing innovative ... Site Reliability Engineer II The SRE II sits at the intersection of software engineering and ...

Site Reliability Engineer I

Sunrise, FL · On-site

$54.25 - $72.25/hr

Experience managing and troubleshooting technology infrastructure and services, including servers, networks, and cloud platforms * Knowledge of cloudbased Site Reliability Engineering (SRE) practices ...

New

Showing results 21-40

Site Reliability Engineer Manager information

See Florida salary details

$8

$47

$68

How much do site reliability engineer manager jobs pay per hour?

As of Sep 4, 2026, the average hourly pay for site reliability engineer manager in Florida is $47.63, according to ZipRecruiter salary data. Most workers in this role earn between $40.96 and $54.42 per hour, depending on experience, location, and employer.

What is a site reliability engineer manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What are the key skills and qualifications needed to thrive as a site reliability engineer manager?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.

How does a site reliability engineer manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How much do site reliability engineer managers get paid?

Site Reliability Engineer Managers typically earn between $120,000 and $180,000 annually, depending on experience, location, and company size. They often oversee teams responsible for system reliability, incident response, and infrastructure automation, requiring strong leadership and technical skills.

Is a Site Reliability Engineer Manager a stressful job?

A Site Reliability Engineer Manager role can be stressful due to the responsibility of maintaining system uptime, managing incident responses, and ensuring reliability across complex infrastructure. The job often involves working under pressure, handling outages, and coordinating teams, but it also offers opportunities for problem-solving and leadership. Stress levels vary depending on company size, team structure, and workload management skills.

What are the most commonly searched types of Site Reliability Engineer jobs in Florida?

The most popular types of Site Reliability Engineer jobs in Florida are:

What cities in Florida are hiring for Site Reliability Engineer Manager jobs?

Cities in Florida with the most Site Reliability Engineer Manager job openings:

Infographic showing various Site Reliability Engineer Manager job openings in Florida as of August 2026, with employment types broken down into 87% Full Time, 12% Part Time, and 1% Contract. Highlights an 79% Physical, 3% Hybrid, and 18% Remote job distribution, with an average salary of $99,078 per year, or $47.6 per hour.

Site Reliability Engineer II

NationsBenefits, LLC

Plantation, FL • On-site

$56.50 - $75.25/hr

Full-time

Medical, PTO

Re-posted 9 days ago


NationsBenefits rating

6.6

Company rating: 6.6 out of 10

Based on 16 frontline employees who took The Breakroom Quiz

317th of 500 rated business services


Job description

NationsBenefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare solutions that drive growth, improve outcomes, reduce costs, and bring value to their members.

Through our comprehensive suite of innovative supplemental benefits, fintech payment platforms, and member engagement solutions, we help health plans deliver high-quality benefits to their members that address the social determinants of health and improve member health outcomes and satisfaction.

Our compliance-focused infrastructure, proprietary technology systems, and premier service delivery model allow our health plan partners to deliver high-quality, value-based care to millions of members.

We offer a fulfilling work environment that attracts top talent and encourages all associates to contribute to delivering premier service to internal and external customers alike. Our goal is to transform the healthcare industry for the better! We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India.

Location: Remote (US-Based Candidates Only)
Site Reliability Engineer II (SRE)

Position Overview

We are seeking a Site Reliability Engineer II (SRE) to join our growing Site Reliability Engineering team. In this role, you will help ensure the availability, reliability, and performance of our production platforms by monitoring systems, responding to incidents, troubleshooting infrastructure issues, and driving automation initiatives.

You will collaborate closely with Development, DevSecOps, and Engineering teams to maintain highly available cloud-native applications while supporting mission-critical healthcare and fintech services.

This position is ideal for someone who enjoys solving production challenges, improving operational efficiency, and working in a fast-paced environment.

Key Responsibilities

Incident Management

  • Serve as the first responder for production incidents by identifying, triaging, and resolving issues.
  • Monitor and respond to alerts generated by Datadog and other monitoring platforms.
  • Perform initial root cause analysis and escalate incidents according to defined SLAs.
  • Communicate incident status and resolution updates to internal stakeholders.
  • Partner with senior engineers to resolve complex production issues.

Monitoring & Platform Reliability

  • Continuously monitor application health, infrastructure performance, and system availability.
  • Configure and optimize monitoring dashboards and alert thresholds.
  • Troubleshoot Kubernetes environments, including pod failures, deployment rollbacks, and log analysis.
  • Support containerized applications running in Kubernetes and Docker environments.

Production Support

  • Participate in a weekday "Follow-the-Sun" production support model with global engineering teams.
  • Participate in an on-call rotation for critical production systems as needed.
  • Help maintain high availability and system uptime.

Automation & Continuous Improvement

  • Develop automation scripts and operational tools using one or more of the following:
    • Python
    • PowerShell
    • Bash
    • C#
    • Java
  • Support CI/CD pipeline monitoring and deployment reliability.
  • Contribute to self-healing solutions and automation initiatives to reduce manual operational tasks.

Collaboration

  • Work closely with Software Engineers, DevSecOps, Infrastructure, and Platform teams.
  • Recommend improvements to monitoring, tooling, and operational processes.
  • Collaborate effectively with globally distributed engineering teams.

Documentation & Compliance

  • Maintain accurate documentation for incidents, troubleshooting procedures, and post-incident reviews.
  • Ensure operational processes align with industry security and compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.

Required Qualifications

  • 3-5 years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Production Support.
  • Hands-on experience with production incident response, troubleshooting, and escalation.
  • Experience with Datadog or similar monitoring and observability platforms.
  • Strong experience with Kubernetes, including monitoring, troubleshooting, and workload management.
  • Experience with Docker or other container technologies.
  • Working knowledge of SQL, MySQL, or NoSQL databases.
  • Ability to work effectively in high-volume, mission-critical production environments.
  • Strong analytical, troubleshooting, and problem-solving skills.
  • Excellent written and verbal communication skills.
  • Willingness to work weekday shifts as part of a global Follow-the-Sun support model.

Preferred Qualifications

  • Experience with cloud platforms such as Microsoft Azure, AWS, or Google Cloud Platform (GCP).
  • Familiarity with CI/CD pipelines and deployment automation.
  • Experience with Helm Charts and Kubernetes deployments.
  • Knowledge of ITIL principles and Agile methodologies.
  • Experience supporting regulated environments such as healthcare or fintech.
  • Scripting or programming experience in Python, PowerShell, Bash, Java, or C#.

Why Join NationsBenefits?

  • Work on technology that positively impacts millions of healthcare members.
  • Join a collaborative, innovative, and supportive engineering culture.
  • Exposure to modern cloud-native technologies and enterprise-scale infrastructure.
  • Competitive compensation and comprehensive benefits.
  • Unlimited Paid Time Off (PTO).
  • Opportunities for career growth and professional development.
  • Work with talented global engineering teams on challenging, high-impact projects.
  • Maintain a healthy work-life balance while contributing to mission-critical platforms.

NationsBenefits is an Equal Opportunity Employer.

Employment Type: FULL_TIME

What NationsBenefits employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom