1

Reliability Manager Jobs in Atlanta, GA (NOW HIRING)

Site Reliability engineer (SRE)

Atlanta, GA · On-site

$54.75 - $72.75/hr

Site Reliability engineer(SRE) Location: Atlanta, GA ( Hybrid - 3days Office - 2 days WFH) Duration ... We specialize in Big Data & Analytics, Digital Transformation, IT Service Management, Cognitive ...

Engineer, Site Reliability

Atlanta, GA · On-site

$84K - $153K/yr

Manage incident response to ensure rapid recovery and minimize service disruption * Adapt to new ... System Reliability (Required) Licenses and Certifications : * Certified Kubernetes Administrator ...

Staff Reliability Engineer

Atlanta, GA · On-site

$165 - $218/hr

ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ... Experience with risk management, change control/change management reviews, and software/firmware ...

Site Reliability Engineer

Atlanta, GA · On-site

$152K - $162K/yr

Manage CI/CD pipeline reliability and deployment quality controls * Conduct root-cause analysis and drive long-term corrective actions * Collaborate with Run teams to transition monitoring ...

Staff Reliability Engineer

Atlanta, GA · On-site

$98K - $124K/yr

ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ... Experience with risk management, change control/change management reviews, and software/firmware ...

Staff Reliability Engineer

Atlanta, GA · On-site

$98K - $124K/yr

ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ... Experience with risk management, change control/change management reviews, and software/firmware ...

Site Reliability Engineer

Atlanta, GA · On-site

$152K - $162K/yr

Manage CI/CD pipeline reliability and deployment quality controls * Conduct root-cause analysis and drive long-term corrective actions * Collaborate with Run teams to transition monitoring ...

Site Reliability Engineer - SRE

Atlanta, GA · On-site

$54.25 - $72/hr

Site Reliability Engineer * Location: Atlanta, GA OR Dallas OR Austin, TX * Duration: Long Term or 6+ Months contract to Hire Note: Remote Possible, however candidates will move to work onsite/Hybrid ...

Site Reliability Engineer

Atlanta, GA · On-site

$155K - $222K/yr

We are one of several SRE teams working together to support a platform that serves more than 500,000 customers and manages over 18 million devices worldwide. The team operates with a high degree of ...

Site Reliability Engineer - SRE

Atlanta, GA · On-site +1

$54.25 - $72/hr

Site Reliability Engineer * Location: Atlanta, GA OR Dallas OR Austin, TX * Duration: Long Term or 6+ Months contract to Hire Note: Remote Possible, however candidates will move to work onsite/Hybrid ...

Site Reliability Engineer

Atlanta, GA · On-site

$100 - $120/hr

Strong knowledge of SRE best practices and incident management protocols * Deep experience using and/or configuring New Relic, Data Dog, SumoLogic or similar observability tools * Proficiency in ...

Manage Customer Reliability Engineering activities driving Application Monitoring, Metrics, Incident Reviews and Long Term Actions, and support BISO activities / implementing InfoSec changes ...

Site Reliability Engineer

Atlanta, GA · On-site +1

$100K - $120K/yr

Strong knowledge of SRE best practices and incident management protocols * Deep experience using and/or configuring New Relic, Data Dog, SumoLogic or similar observability tools * Proficiency in ...

Showing results 41-60

Reliability Manager information

See Atlanta, GA salary details

$59.6K

$113K

$162K

How much do reliability manager jobs pay per year?

As of Aug 11, 2026, the average yearly pay for reliability manager in Atlanta, GA is $112,983.00, according to ZipRecruiter salary data. Most workers in this role earn between $90,900.00 and $134,600.00 per year, depending on experience, location, and employer.

What is the role of a reliability manager?

A reliability manager oversees the maintenance and reliability of equipment and systems within an organization to ensure optimal performance and minimize downtime. They analyze failure data, develop maintenance strategies, and implement continuous improvement processes, often using tools like root cause analysis and reliability-centered maintenance. Strong analytical skills, technical knowledge, and certifications such as Certified Reliability Engineer (CRE) are typically required for this role.

Is reliability engineering in demand?

Reliability engineering is in high demand across industries such as manufacturing, energy, and aerospace, as companies prioritize system uptime and maintenance efficiency. Reliability Managers with skills in data analysis, failure modes, and certifications like RCM are sought after to improve equipment performance and reduce downtime.

What are the key skills and qualifications needed to thrive as a reliability manager?

A Reliability Manager needs strong analytical skills, a solid background in engineering or maintenance, and experience with reliability-centered maintenance methodologies. Familiarity with tools like Failure Mode and Effects Analysis (FMEA), Root Cause Analysis (RCA), and certifications such as Certified Reliability Engineer (CRE) are often required. Leadership, problem-solving, and the ability to communicate complex technical information clearly are crucial soft skills for this role. These skills help ensure equipment uptime, optimize maintenance processes, and foster a culture of continuous improvement within the organization.

What does a reliability manager do?

A Reliability Manager is responsible for ensuring that equipment, processes, and systems operate efficiently and consistently to minimize downtime and maximize performance. They develop and implement reliability strategies, conduct root cause analyses, and oversee preventive and predictive maintenance programs. Their role involves working closely with maintenance teams, engineers, and production staff to improve asset reliability and extend equipment lifespan. Additionally, they analyze failure data, recommend improvements, and help optimize operational costs through reliability-centered maintenance practices.

What are the most commonly searched types of Reliability jobs in Atlanta, GA? The most popular types of Reliability jobs in Atlanta, GA are:
What are popular job titles related to Reliability Manager jobs in Atlanta, GA? For Reliability Manager jobs in Atlanta, GA, the most frequently searched job titles are:
What job categories do people searching Reliability Manager jobs in Atlanta, GA look for? The top searched job categories for Reliability Manager jobs in Atlanta, GA are:
What cities near Atlanta, GA are hiring for Reliability Manager jobs? Cities near Atlanta, GA with the most Reliability Manager job openings:
Infographic showing various Reliability Manager job openings in Atlanta, GA as of August 2026, with employment types broken down into 87% Full Time, and 13% Part Time. Highlights an 93% In-person, and 7% Remote job distribution, with an average salary of $112,983 per year, or $54.3 per hour.

Site Reliability Engineer

PVH (Tommy Hilfiger/Calvin Klein)

Alpharetta, GA • On-site

$120 - $180/hr

Other

Posted 5 days ago


PVH Corp. rating

6.3

Company rating: 6.3 out of 10

Based on 7 frontline employees who took The Breakroom Quiz


Job description

Company Overview:

About Us: Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 districts. Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 million students and educators nationwide. By connecting technology and operational workflows, Incident IQ enables schools to streamline processes, reduce administrative burdens, and focus on what matters most: supporting students.

Purpose: Incident IQ is committed to creating a future where every K-12 district operates with seamless efficiency. When operations are unified on a single platform, districts gain the clarity and control needed to build a stronger foundation for student success. We're focused on delivering the tools, support, and partnerships that help make that vision a reality.

Mission: Incident IQ is on a mission to eliminate the friction of disconnected systems and clunky workflows that slow schools down. We're reimagining the critical work that happens behind the scenes, bringing visibility, efficiency, and impact to the processes that keep classrooms running. By streamlining the complex, automating the routine, and surfacing the insights that matter most, we can create the conditions for educators to teach, students to thrive, and districts to shape the future of education.

Site Reliability Engineer (SRE) Overview:

We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first dedicated Site Reliability Engineer, and you'll be defining what "reliable" means for our production systems, not maintaining someone else's playbook. You'll work with leading-edge observability and reliability tooling, and the calls you make will directly shape how confidently the whole engineering org ships.

Expect real engineering deep dives, not top-down mandates. We love digging into a hard problem together, and we want you to bring a strong point of view, back it up with data and sound reasoning, and enjoy the back-and-forth as we work toward the best answer. Good persuasion skills matter here as much as technical depth, since good ideas still have to win the room. We move at startup speed: we'd rather figure something out in a few hours than plan it for weeks. We're a collaborative, respectful team: we debate ideas hard, never people.

We care much more about a proven track record running big, ambiguous projects efficiently than about years of tenure or a wall of certifications. You should be genuinely comfortable working independently: we won't hand-hold you or chase you for status updates. We expect you to take total ownership of outcomes and drive them without being asked twice, and without running your own separate agenda. This work is relentless, juggling several things at once under real time pressure is normal here, and the right candidate is passionate about SRE and thrives on that intensity, not just tolerates it.

Site Reliability Engineer (SRE) Responsibilities:
  • SLI/SLO Definition & Grafana Implementation: Drive the definition of Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for our core services, translating them into insightful Grafana dashboards and actionable, burn-rate-based alerting, so pages are precise and noise stays low.
  • Incident Management: Stand up our incident management practice (tooling such as PagerDuty, on-call training, incident command), then own and continuously improve it, stepping in personally only for the most severe incidents.
  • Observability Stack Ownership: Own the observability stack end to end: metrics, logs, traces, Real User Monitoring (RUM), and synthetic checks across the user journey, alerting whenever a signal deviates from baseline.
  • Team Enablement: Partner with engineering teams to refine SLIs, SLOs, and error budgets as services evolve, and coach teams on SRE and observability best practices.
  • Toil Reduction: Identify and automate away manual, repetitive operational work through infrastructure as code and tooling.
  • Chaos & Performance Engineering: Design and run load/performance tests and chaos engineering game days to proactively surface weaknesses before they cause incidents.

For example: in your first few days, you might stand up an SLO and a burn-rate alert in Grafana for our highest-traffic service. Within a couple of weeks, PagerDuty on-call is configured and the rotation is trained on incident command. That's the pace we operate at here: hours and days, not weeks.

Site Reliability Engineer (SRE) Requirements:
  • Education & Systems Foundations: Bachelor's degree in Computer Science, Computer Engineering, or equivalent f
#J-18808-Ljbffr

What PVH Corp. employees say

Pay

Hours and flexibility

Workplace

Get the full story on Breakroom