1

Power Plant Reliability Engineer Jobs in Georgia

Reliability Engineer

Tucker, GA ยท On-site

$96K - $121K/yr

Job Purpose Provide maintenance reliability systems and equipment consulting for the facility ... Partner with the Maintenance Manager and Plant Director to identify and recommend capital ...

New

Site Reliability Engineer (SRE)

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Leverage AI-powered engineering tools and agentic AI capabilities to improve monitoring, incident ... SRE, DevOps, Production Support, or Site Reliability Engineering role. * Strong experience ...

In this role, the engineer will ensure the reliability, scalability, in an Azure cloud environment ... That's the power of true partnership. TEKsystems is an Allegis Group company. The company is an ...

New

Reliability Engineer

Atlanta, GA ยท On-site

$68 - $75/hr

In this role, the engineer will ensure the reliability, scalability, in an Azure cloud environment ... That's the power of true partnership. TEKsystems is an Allegis Group company. The company is an ...

New

Equipment Reliability Engineer

Perry, GA ยท On-site

$92K - $116K/yr

... all plant assets. Identify: Critical Assets Equipment impacting: * Food safety * Throughput ... Power BI * Vibration Analysis * Thermography Additional Information The starting salarie for this ...

Equipment Reliability Engineer

Perry, GA ยท On-site

$92K - $116K/yr

... all plant assets. Identify: Critical Assets Equipment impacting: * Food safety * Throughput ... Power BI * Vibration Analysis * Thermography Additional Information The starting salarie for this ...

Equipment Reliability Engineer

Perry, GA ยท On-site

$92K - $116K/yr

... all plant assets. Identify: Critical Assets Equipment impacting: * Food safety * Throughput ... Power BI * Vibration Analysis * Thermography Additional Information The starting salarie for this ...

Site Reliability Engineer

Norcross, GA ยท On-site

$53.50 - $71/hr

Powering the world's payments ecosystem ACI powers the payments ecosystem - globally, and you power ... As a Senior Site Reliability Engineer (SRE) - Azure & GitOps (CI/CD) in Norcross, GA or Omaha, NE ...

Site Reliability Engineer

Norcross, GA ยท Hybrid

$53.50 - $71/hr

Powering the world's payments ecosystem ACI powers the payments ecosystem - globally, and you power ... As a Senior Site Reliability Engineer (SRE) - Azure & GitOps (CI/CD) in Norcross, GA or Omaha, NE ...

Site Reliability Engineer

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 ... Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to ...

Site Reliability Engineer

Alpharetta, GA ยท On-site

$55.75 - $74/hr

Trusted by over 2,000 districts, Incident IQ powers mission-critical services for more than 12 ... Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to ...

Reliability Engineer (22895)

Gainesville, GA ยท On-site

$95K - $120K/yr

The Reliability Engineer is responsible for ensuring the reliability, maintainability, and ... plant operations. Essential Job Responsibilities: * Development and refinement of the preventive ...

Staff Reliability Engineer

Atlanta, GA ยท On-site

$98K - $124K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

The Plant Engineering Manager oversees all engineering, maintenance, utilities, and automation ... Reliability Engineering - Lead root cause analysis (RCA), FMEA (FMECA, FMEDA), and ...

Reliability Engineer II

Savannah, GA ยท On-site

$95K - $120K/yr

The Engineer II will also focus on optimizing equipment maintainability, extending equipment life ... years of plant maintenance experience * 2+years of experience in maintenance and reliability ...

Showing results 21-40

Power Plant Reliability Engineer information

What does a power plant reliability engineer do?

A Power Plant Reliability Engineer is responsible for ensuring the reliable and efficient operation of power plant equipment and systems. They analyze equipment performance, investigate failures, and develop strategies to prevent future issues. Their work often involves monitoring maintenance practices, recommending improvements, and implementing reliability-centered maintenance programs. By optimizing plant reliability, they help reduce downtime, lower maintenance costs, and ensure consistent power generation.

What are the key skills and qualifications needed to thrive as a power plant reliability engineer?

To thrive as a Power Plant Reliability Engineer, you need a solid background in mechanical, electrical, or chemical engineering, typically with a relevant degree and experience in power generation environments. Familiarity with reliability-centered maintenance (RCM) methodologies, root cause analysis tools, and computerized maintenance management systems (CMMS) is essential. Strong analytical thinking, problem-solving abilities, and effective communication skills help you collaborate across teams and drive continuous improvement. These skills ensure plant equipment operates safely and efficiently, minimizing downtime and maximizing operational reliability.

What are the most common challenges a power plant reliability engineer faces when implementing reliability improvement initiatives?

Power Plant Reliability Engineers often encounter challenges such as resistance to change from operations and maintenance teams, limited availability of historical failure data, and balancing short-term production goals with long-term reliability improvements. They must collaborate closely with cross-functional teams to analyze equipment performance, identify root causes of failures, and develop actionable maintenance strategies. Building effective communication channels and fostering a culture of reliability are key to overcoming these challenges and ensuring ongoing plant performance.

What is the difference between Power Plant Reliability Engineer vs Power Plant Maintenance Engineer?

AspectPower Plant Reliability EngineerPower Plant Maintenance Engineer
CredentialsEngineering degree, reliability certificationsTechnical diploma or engineering background
Work EnvironmentAnalysis, planning, and optimization in office and plant settingsHands-on maintenance and repair in plant facilities
Primary FocusEnsuring equipment reliability and reducing downtimePerforming maintenance and repairs to keep equipment operational
Industry UsageUsed across power generation companies for reliability strategiesCommonly employed in plant operations and maintenance teams

The Power Plant Reliability Engineer focuses on analyzing and improving equipment reliability through data analysis and strategic planning, while the Power Plant Maintenance Engineer handles the hands-on maintenance and repairs. Both roles are essential for efficient power plant operation but differ in their primary responsibilities and work environment.

Are power plant reliability engineers in demand?

Power plant reliability engineers are in steady demand due to the need for maintaining and optimizing energy generation facilities. Their expertise in equipment maintenance, troubleshooting, and reliability analysis is essential for ensuring continuous power supply, especially as the energy sector emphasizes safety and efficiency. Job opportunities often require knowledge of industry standards and certifications such as NERC or API.

What are popular job titles related to Power Plant Reliability Engineer jobs in Georgia?

For Power Plant Reliability Engineer jobs in Georgia, the most frequently searched job titles are:

What job categories do people searching Power Plant Reliability Engineer jobs in Georgia look for?

The top searched job categories for Power Plant Reliability Engineer jobs in Georgia are:

What cities in Georgia are hiring for Power Plant Reliability Engineer jobs?

Cities in Georgia with the most Power Plant Reliability Engineer job openings:

Infographic showing various Power Plant Reliability Engineer job openings in Georgia as of August 2026, with employment types broken down into 100% Contract. Highlights an 100% In-person job distribution.

Site Reliability Engineer (SRE)

Atlanta, GA โ€ข On-site

Atlanticus Holdings Corporation
Finance and Insuranceย โ€ขย 201 - 500 employees

$100K - $115K/yr

Other

Medical, Retirement, PTO

Re-posted 6 days ago


Key responsibilities

  • Provide 24x7 production support through an on-call rotation to ensure application availability and rapid incident response.

  • Continuously monitor production applications, infrastructure, and platform health using Datadog, Splunk, CloudWatch, and other monitoring tools.

  • Respond to production incidents, troubleshoot issues, and restore services while minimizing customer impact.


Job description

Site Reliability Engineer (SRE)
When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage entrepreneurial thinking to empower our customers toward financial well-being.
Atlanticusโ„ข technology enables bank, retail, and healthcare partners to offer more inclusive financial services to everyday Americans through the use of proprietary analytics. We apply the experience gained and infrastructure built from servicing over 20 million customers and over $40 billion in consumer loans over more than 25 years of operating history to support lenders that originate a range of consumer loan products. These products include retail and healthcare, private label credit and general-purpose credit cards marketed through our omnichannel platform, including retail point-of-sale, healthcare point-of-care, direct mail solicitation, digital marketing, and partnerships with third parties. Additionally, through our Auto Finance subsidiary, Atlanticus serves the individual needs of automotive dealers and automotive non-prime financial organizations with multiple financing and service programs.
Office Locations available for this role include
  • Austin, TX - Situated in The Domain, a vibrant tech hub with park-like surroundings, top restaurants, and convenient parking, perfect for post-work socializing.
  • Atlanta, GA - Located in the Queen Building (King & Queen Towers, Sandy Springs), with easy access to I-285, GA-400, and a free shuttle to MARTA.

Work Culture
We foster a collaborative, innovative environment where everyone contributes to building something meaningful. You'll be empowered to lead, grow, and make an impact.
The Role
We are seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational excellence of our cloud-native applications running on AWS. This is a hands-on role responsible for monitoring and supporting production systems, automating operational tasks, managing deployments, and driving continuous improvements in system stability.
The ideal candidate has strong experience supporting Java-based applications running on Amazon EKS, a solid understanding of AWS infrastructure, and expertise with observability platforms such as Datadog and Splunk. This role requires participation in a 24x7 production support and on-call rotation, working closely with Development, DevOps, IT Ops, Database, Network, and Security teams to maintain highly available production services.
The successful candidate should be passionate about automation, troubleshooting complex production issues, improving application reliability, and leveraging AI-powered tools to enhance operational efficiency.
Key Responsibilities
  • Provide 24x7 production support through an on-call rotation to ensure application availability and rapid incident response.
  • Continuously monitor production applications, infrastructure, and platform health using Datadog, Splunk, CloudWatch, and other monitoring tools.
  • Respond to production incidents, troubleshoot issues, and restore services while minimizing customer impact.
  • Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
  • Deploy and support Java-based applications running on Docker and Amazon EKS using CI/CD pipelines.
  • Execute production deployments, application releases, hotfixes, and rollbacks following change management processes.
  • Monitor and manage scheduled application jobs, batch processes, and integrations to ensure successful execution.
  • Troubleshoot Java application issues using logs, JVM metrics, thread dumps, heap dumps, and application performance metrics.
  • Analyze application, infrastructure, and Kubernetes logs using Splunk and Datadog to identify performance bottlenecks and operational issues.
  • Develop automation scripts using Python, Bash, or similar scripting languages to eliminate repetitive operational tasks.
  • Build self-healing and automated operational processes to improve system reliability and reduce manual intervention.
  • Support Kubernetes (Amazon EKS) environments, including troubleshooting pods, deployments, networking, ingress, and scaling issues.
  • Maintain and improve dashboards, alerts, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational runbooks.
  • Partner with Development teams to improve application reliability, resiliency, scalability, and performance.
  • Continuously improve operational processes, monitoring coverage, automation, and deployment practices.
  • Leverage AI-powered engineering tools and agentic AI capabilities to improve monitoring, incident response, automation, and operational efficiency.

You're a great fit if you have
  • 5+ years of experience supporting production applications in an SRE, DevOps, Production Support, or Site Reliability Engineering role.
  • Strong experience supporting Java-based applications in production environments.
  • Hands-on experience with AWS services including EKS, EC2, ALB/NLB, RDS, IAM, Route 53, CloudWatch, S3, and VPC.
  • Experience with Kubernetes (Amazon EKS), Docker, and containerized application deployments.
  • Strong experience using Datadog / Splunk for infrastructure monitoring, APM, troubleshooting, dashboards, alerting, and log analysis.
  • Experience performing production deployments through CI/CD pipelines (Jenkins, GitHub Actions, Argo CD, or similar).
  • Experience supporting MySQL and Oracle databases from an application support perspective.
  • Proficiency in Python, Bash, or other scripting languages for automation.
  • Strong Linux system administration and troubleshooting skills.
  • Excellent troubleshooting skills across distributed applications, networking, and cloud infrastructure.
  • Knowledge of networking fundamentals including DNS, TCP/IP, HTTP/HTTPS, TLS, load balancing, and firewalls.
  • Experience with incident management, problem management, and change management processes.
  • Experience using AI-assisted development tools or agentic AI systems to improve operational efficiency.

Preferred
  • Experience with Helm and GitOps deployment models.
  • Experience with Terraform or Infrastructure as Code.
  • Familiarity with Prometheus, Grafana, or OpenTelemetry.
  • Experience supporting microservices architectures.
  • Knowledge of JVM tuning and Java performance optimization.
  • Experience with AWS Auto Scaling, Karpenter, or Cluster Autoscaler.

Why You'll Love Working Here
This isn't just a job, it's a place to lead, grow, and thrive. If you believe in your skills and drive, we'll provide the resources and support to help you succeed.
Benefits include
  • Generous PTO and holiday schedule
  • 401(k) with company match
  • Employee stock purchase plan
  • Ongoing training (lunch & learns, financial and health webinars)
  • Team volunteer outings

Atlanticus is an equal opportunity employer. All qualified applicants will receive consideration without regard to race, religion, gender, sexual orientation, age, veteran status, disability, or other protected status.