1

Reliability Engineer Manager Jobs in Canton, GA (NOW HIRING)

GCP SRE

Alpharetta, GA · On-site

$55.75 - $74/hr

This role is for a Site Reliability Engineer (SRE) with a strong emphasis on Google Cloud Platform ... The position requires robust incident management capabilities within the GCP ecosystem, focusing on ...

Senior Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

Who We Are QGenda is redefining healthcare workforce management everywhere care is delivered. We're ... About Your Role As a Senior Site Reliability Engineer, you will work with our Infrastructure and ...

Database SRE

Alpharetta, GA · On-site

$55.75 - $74/hr

Database SRE Location: Alpharetta, Georgia Duration: Contract Job ID: 179116 Job Overview: We are ... Performance & Incident Management: * Analyze and resolve database performance issues.

New

Senior Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

Design, implement, and manage scalable systems that ensure high availability, fault tolerance, and ... Actively contribute to fostering an SRE culture within the organization by promoting observability ...

Site Reliability Engineer (AWS)

Atlanta, GA · Hybrid

$54.75 - $72.75/hr

This position is under our CTO org to support SRE functions for innovation and growth for the ... Manage deployment pipelines and configuration management for consistent and reliable app ...

Site Reliability Engineer

Atlanta, GA · On-site

$152K - $162K/yr

Unum Group seeks Site Reliability Engineers in Atlanta, GA. Applicants who are interested in this ... Manage CI/CD pipeline reliability and deployment quality controls * Conduct root-cause analysis and ...

Site Reliability Engineer (AWS)

Atlanta, GA · Hybrid

$54.75 - $72.75/hr

This position is under our CTO org to support SRE functions for innovation and growth for the ... Manage deployment pipelines and configuration management for consistent and reliable app ...

Site Reliability Engineer - Networking

Atlanta, GA · On-site

$54.75 - $72.75/hr

As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and ... Design, implementation and management of an overlay network to support 1000's of containers.

Showing results 41-60

Reliability Engineer Manager information

See Canton, GA salary details

$57.6K

$111.4K

$133.1K

How much do reliability engineer manager jobs pay per year?

As of Aug 23, 2026, the average yearly pay for reliability engineer manager in Canton, GA is $111,388.00, according to ZipRecruiter salary data. Most workers in this role earn between $96,800.00 and $121,800.00 per year, depending on experience, location, and employer.

What does a reliability engineer manager do?

A Reliability Engineer Manager oversees teams responsible for improving the reliability and performance of systems, machinery, or processes within an organization. They develop maintenance strategies, lead root cause analyses of failures, and implement best practices to minimize downtime and costs. Additionally, they collaborate with other departments to ensure that reliability goals align with business objectives and compliance standards. Their role is crucial in industries such as manufacturing, energy, and technology, where system uptime and safety are critical.

What are the key skills and qualifications needed to thrive as a reliability engineer manager?

To thrive as a Reliability Engineer Manager, you need a strong background in engineering principles, reliability analysis, and maintenance strategies, typically supported by a degree in engineering and experience in reliability roles. Familiarity with reliability-centered maintenance (RCM), failure mode and effects analysis (FMEA), and asset management software such as SAP or Maximo is common, along with certifications like Certified Reliability Engineer (CRE). Leadership, problem-solving, and effective communication are vital soft skills for managing teams and driving cross-functional initiatives. These competencies are crucial for minimizing downtime, optimizing equipment performance, and ensuring long-term operational efficiency.

What are some common challenges reliability engineer managers face when balancing long-term reliability improvements with immediate operational demands?

Reliability Engineer Managers often need to prioritize urgent maintenance issues while also driving long-term reliability initiatives. Balancing these competing demands can be challenging, as immediate equipment failures may require quick fixes that temporarily interrupt ongoing improvement projects. Effective managers work closely with operations, maintenance, and engineering teams to communicate priorities, allocate resources, and implement sustainable solutions that address root causes rather than just symptoms. This role typically involves using data-driven decision-making and fostering a culture of proactive maintenance and continuous improvement.

What is the difference between Reliability Engineer Manager vs Reliability Engineer?

AspectReliability EngineerReliability Engineer Manager
Required CredentialsBachelor's in Engineering or related field; certifications like CRC, CRESame as Reliability Engineer, plus leadership experience
Work EnvironmentDesign, analyze, and improve system reliability; often in teamsOversees Reliability Engineers; manages projects and teams
Employer & Industry UsageManufacturing, aerospace, energy, automotiveSame industries, with added managerial responsibilities
Common Search & ComparisonFocuses on technical skills and hands-on reliability tasksFocuses on leadership, team management, and strategic planning

The main difference between a Reliability Engineer and a Reliability Engineer Manager lies in their responsibilities. The Reliability Engineer focuses on technical analysis and system improvements, while the Reliability Engineer Manager oversees teams, manages projects, and develops strategies to enhance reliability across the organization.

Infographic showing various Reliability Engineer Manager job openings in Canton, GA as of August 2026, with employment types broken down into 45% Full Time, and 55% Contract. Highlights an 91% In-person, and 9% Remote job distribution, with an average salary of $111,388 per year, or $53.6 per hour.

Senior Site Reliability Engineer

IRB USA Inspire Resources

Atlanta, GA • On-site

$130 - $180/hr

Other

Posted 3 days ago

New


Job description

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale. The ideal candidate has hands‑on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.

RESPONSIBILITIESReliability Engineering
  • Define and manage SLIs, SLOs, and Error Budgets for critical services
  • Drive production readiness reviews and reliability requirements into architecture and design
  • Perform capacity planning, failure mode analysis, and dependency risk assessments
  • Identify systemic reliability risks and drive remediation before they cause customer impact
Observability
  • Design monitoring, alerting, logging, and tracing solutions using modern observability tooling
  • Improve signal-to-noise ratio and reduce alert fatigue
  • Build dashboards and telemetry that reflect true service health, not just infrastructure metrics
Incident Management
  • Lead technical response for high-severity incidents
  • Drive blameless postmortems and root cause analysis focused on systemic fixes
  • Continuously improve detection, response, and recovery processes
  • Participate in an on‑call rotation
Automation & Toil Reduction
  • Identify and eliminate manual, repetitive operational work through automation
  • Build self‑healing systems, tooling, and scripts to reduce human intervention
  • Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green)
  • Support Infrastructure as Code (Terraform, Bicep, or similar)
Performance & Scalability
  • Conduct load testing, performance benchmarking, and bottleneck analysis
  • Partner with engineering to design systems for horizontal scalability and fault tolerance
Collaboration & Culture
  • Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)
  • Mentor engineers on SRE best practices
  • Promote a culture of engineering‑driven reliability over reactive operations
EDUCATION AND EXPERIENCE QUALIFICATIONSRequired Qualifications
  • 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
  • 2+ years experience with Kubernetes and containerized workloads
  • 4-year degree in Computer Science or related field
Preferred Qualifications
  • Experience with chaos engineering or resiliency testing
  • Experience with high-volume, high-availability transactional systems
  • Experience with AI-assisted observability or operational automation
  • Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms
REQUIRED KNOWLEDGE, SKILLS, OR ABILITIES
  • Strong programming/scripting skills (Python, Go, Java, or Node.js)
  • Demonstrated experience defining and operating against SLOs/Error Budgets
  • Strong skills in leading incident response and root cause analysis for production systems
  • Solid understanding of distributed systems and microservices architecture
  • Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)
  • Expertise with observability platforms and monitoring strategy

This position is based in our Atlanta Support Center, with an expected on‑site presence of 80%.

Inspire is a multi‑brand restaurant company whose portfolio includes more than 33,300 Arby’s, Baskin‑Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide. We’re made up of some of the world’s most iconic restaurant brands, but we’re much more than just a restaurant company. We’re a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple—it’s an experience. At Inspire, that’s our purpose: to ignite and nourish flavorful experiences. Inspire is a multi‑brand restaurant company whose portfolio includes more than 33,400 Arby’s, Baskin‑Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants across 55 global markets. In an industry facing increasing disruption, our leaders saw an opportunity to build a restaurant company unlike any other – one that brings together differentiated yet complementary brands and aims to make them stronger than they would be on their own. Found inherently in the purposes of our family of brands, we identified a common thread between our restaurants – the capacity to inspire. From guest experience to career development to community well‑being, Inspire plays a role in the lives of millions of people every day. Our brands are diverse, distinctive, and fan favorites. In a sense, you could say we seek those who provide something different than the norm.

#J-18808-Ljbffr