1

Sr Reliability Engineer Jobs in Georgia (NOW HIRING)

Site Reliability Engineer

Alpharetta, GA ยท On-site

$65 - $75/hr

Title: Senior Site Reliability Engineer Location: Alpharetta, GA Duration: 6-12+ Months About the Role We're seeking an experienced Senior Site Reliability Engineer to join our team and play a ...

Sr. Site Reliability Engineer

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months Contract Visa : US Citizens/ Green Card Need Local to Atlanta Metro Area for F2F interview Notes from ...

Falcomm is seeking an RFIC Reliability Engineer to lead reliability analysis and qualification activities for RF integrated circuits and semiconductor products. This role will focus on evaluating ...

Senior Site Reliability Engineer II

Alpharetta, GA ยท On-site +1

$125K - $209K/yr

We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role you will be ...

Senior Site Reliability Engineer II

Buford, GA ยท On-site +1

$125K - $209K/yr

We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role you will be ...

Senior Site Reliability Engineer II

Atlanta, GA ยท On-site +1

$125K - $209K/yr

We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role you will be ...

Falcomm is seeking an RFIC Reliability Engineer to lead reliability analysis and qualification activities for RF integrated circuits and semiconductor products. This role will focus on evaluating ...

Falcomm is seeking an RFIC Reliability Engineer to lead reliability analysis and qualification activities for RF integrated circuits and semiconductor products. This role will focus on evaluating ...

SRE/DevOps Engineer

Johns Creek, GA ยท On-site

$52.75 - $70.25/hr

Collaborate with senior engineers to improve deployment reliability, scalability, and operational efficiency. * Utilize AI-powered developer tools (such as Claude) to improve engineering productivity ...

SRE Lead/ Architect

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Job Title: SRE Lead/Architect Location: Atlanta, GA - Hybrid (Thur to next wed (Alternate weeks ... Act as a senior technical advisor and subject matter expert on reliability, scalability, and ...

Showing results 21-40

Sr Reliability Engineer information

See Georgia salary details

$18

$54

$77

How much do sr reliability engineer jobs pay per hour?

As of Aug 13, 2026, the average hourly pay for sr reliability engineer in Georgia is $54.39, according to ZipRecruiter salary data. Most workers in this role earn between $44.86 and $65.14 per hour, depending on experience, location, and employer.

What is the difference between Sr Reliability Engineer vs Reliability Engineer?

AspectSr Reliability EngineerReliability Engineer
CredentialsBachelor's or higher in engineering, certifications like CRE or Six Sigma often preferredBachelor's degree in engineering or related field, similar certifications
Work EnvironmentTypically in manufacturing, energy, or tech industries focusing on system reliability and failure analysisSimilar industries, focusing on designing, testing, and improving product or system reliability
Employer UsageUsed in companies seeking experienced engineers to lead reliability projectsUsed for entry to mid-level roles focused on reliability assessments

The main difference is experience level and responsibility. Sr Reliability Engineers often lead projects and have more advanced certifications, while Reliability Engineers focus on supporting reliability tasks. Both roles require similar credentials and work in comparable environments, but the senior role involves more leadership and strategic planning.

How does a Sr Reliability Engineer typically collaborate with cross-functional teams to improve system reliability?

As a Sr Reliability Engineer, you will regularly work alongside operations, development, and QA teams to identify potential reliability risks and implement solutions. This often involves facilitating root cause analyses after incidents, sharing best practices, and leading reliability-focused design reviews. Effective communication and the ability to translate complex technical findings into actionable recommendations are key, as your insights directly influence infrastructure and product decisions. This collaborative approach helps foster a culture of reliability across the organization.

What are the key skills and qualifications needed to thrive as a Sr Reliability Engineer?

To thrive as a Sr Reliability Engineer, you need expertise in reliability engineering principles, root cause analysis, and a relevant engineering degree, often with several years of industry experience. Familiarity with reliability software (such as ReliaSoft or Minitab), maintenance management systems, and certifications like Certified Reliability Engineer (CRE) are commonly required. Strong problem-solving abilities, proactive communication, and leadership skills help you drive reliability initiatives and collaborate across teams. These competencies are vital to ensure equipment uptime, reduce failures, and improve operational efficiency in complex industrial environments.

What does a Sr Reliability Engineer do?

A Sr Reliability Engineer is responsible for ensuring that products, systems, or processes operate reliably and efficiently over their expected lifecycle. They analyze failure data, develop reliability test plans, and implement strategies to predict and prevent failures. Their role often involves collaborating with design, manufacturing, and maintenance teams to improve product quality and reduce downtime. Additionally, they may use reliability modeling tools and statistical techniques to assess risk and recommend improvements. This position typically requires advanced engineering knowledge and experience in reliability engineering principles.

What cities in Georgia are hiring for Sr Reliability Engineer jobs?

Cities in Georgia with the most Sr Reliability Engineer job openings:

Infographic showing various Sr Reliability Engineer job openings in Georgia as of August 2026, with employment types broken down into 87% Full Time, 6% Part Time, 1% Temporary, and 6% Contract. Highlights an 86% Physical, 5% Hybrid, and 9% Remote job distribution, with an average salary of $113,131 per year, or $54.4 per hour.

Site Reliability Engineer

Amicis Global

Alpharetta, GA โ€ข On-site

$65 - $75/hr

Contractor

Medical, Dental, Vision

Re-posted 26 days ago


Job description

Title: Senior Site Reliability Engineer
Location: Alpharetta, GA
Duration: 6-12+ Months
About the Role
We're seeking an experienced Senior Site Reliability Engineer to join our team and play a critical role in ensuring the reliability, scalability, and performance of our cloud infrastructure. You'll be a technical leader who combines deep operational expertise with strong automation skills to build and maintain highly available systems. As a Kubernetes expert, you'll drive our container orchestration strategy and serve as a technical authority for our platform teams.
Key Responsibilities:
Infrastructure & Automation

Design, deploy, and manage cloud infrastructure across AWS and Azure using Terraform and infrastructure-as-code principles
Architect, deploy, and maintain production-grade Kubernetes clusters with a focus on reliability, security, and performance
Serve as the subject matter expert on Kubernetes, providing guidance and best practices to engineering teams
Build and maintain automated provisioning pipelines to ensure consistent, repeatable deployments
Implement and maintain HashiCorp Vault on AWS for secrets management and security, including Vault integration with Kubernetes
Design and implement automated High Availability and Disaster Recovery (HA/DR) capabilities through CI/CD pipelines
Optimize cloud resources and Kubernetes workloads for performance, cost efficiency, and reliability.
Observability & Monitoring
Architect and implement comprehensive observability solutions using Datadog for cloud-native applications and Kubernetes infrastructure
Build monitoring, logging, and alerting frameworks for containerized workloads that provide actionable insights into system health
Implement Kubernetes-native monitoring patterns and troubleshoot complex container orchestration issues
Integrate Datadog with PagerDuty and other incident management platforms
Define and track SLIs, SLOs, and error budgets to drive reliability improvements
Create custom dashboards and monitors to track infrastructure, application, and Kubernetes cluster performance
CI/CD & Pipeline Management
Design, build, and maintain robust CI/CD pipelines that enable rapid, safe deployments to Kubernetes
Implement GitOps workflows and automated deployment strategies for containerized applications
Implement automated testing, security scanning, and quality gates within pipelines
Drive solutions through test, QA, and production environments with appropriate controls and safeguards
Automate deployment strategies including blue-green, canary, and rolling deployments in Kubernetes
Security & Vulnerability Management
Identify, assess, and remediate security vulnerabilities in infrastructure, applications, and Kubernetes clusters
Implement Kubernetes security best practices including RBAC, pod security policies/standards, and network policies
Collaborate with security teams to implement and maintain security best practices
Manage and maintain HashiCorp Vault infrastructure for secure secrets management
Ensure compliance with security policies and industry standards across all environments
Incident Management & Response
Participate in 24/7 on-call rotation to respond to critical production incidents
Serve as Incident Commander, coordinating cross-functional response teams during major outages
Lead post-incident reviews and drive thorough root cause analysis across engineering teams
Troubleshoot complex Kubernetes and distributed systems issues under pressure
Develop and refine incident response procedures and runbooks
Collaboration & Leadership
Partner with engineering teams to improve system reliability and performance
Mentor junior SREs and promote SRE best practices across the organization
Lead Kubernetes adoption efforts and educate teams on container orchestration best practices
Drive initiatives to reduce toil through automation and process improvement
Contribute to architectural decisions with a reliability and operability lens
 
Required Qualifications:
5+ years of experience in Site Reliability Engineering, DevOps, or similar roles
Expert-level knowledge of Kubernetes
, including architecture, operations, and troubleshooting in production environments
Proven track record as a go-to Kubernetes resource and technical authority
Deep understanding of container technologies (Docker, containerd) and orchestration patterns
Strong hands-on experience with
AWS and Azure
cloud platforms
Proficiency in
Terraform
for infrastructure automation and management
Expert-level knowledge of
Datadog
for monitoring, logging, and observability
Experience with
HashiCorp Vault
, including deployment and management on AWS and Kubernetes integration
Deep understanding of
CI/CD pipelines
, including design, implementation, and optimization for containerized workloads
Proven ability to implement automated HA/DR solutions through CI/CD workflows
Strong programming skills in
Python
for automation, tooling, and analysis
Proven experience building observability solutions for distributed cloud applications
Experience configuring monitoring and alerting systems and integrating with paging platforms like PagerDuty
Demonstrated experience identifying and remediating security vulnerabilities
Experience driving deployments through multiple environments (test/QA/production) with proper gates and controls
Demonstrated experience participating in on-call rotations and responding to production incidents
Experience serving as Incident Commander or leading incident response efforts
Track record of conducting root cause analysis and driving systemic improvements
Strong understanding of networking, security, and cloud architecture principles
Excellent communication skills with ability to work across multiple teams and explain complex Kubernetes concepts
 
Preferred Qualifications:
Experience with
Google Cloud Platform (GCP)
and GKE
Certified Kubernetes Administrator (CKA) or Certified Kubernetes Security Specialist (CKS)
Experience with service mesh technologies (Istio, Linkerd, Consul)
Knowledge of Helm, Kustomize, and other Kubernetes tooling
Experience with GitOps tools (ArgoCD, Flux)
Familiarity with additional CI/CD tools (Jenkins, GitLab CI, GitHub Actions, CircleCI)
Experience with configuration management tools (Ansible, Chef, Puppet)
Background in software engineering or systems programming
Understanding of chaos engineering and reliability testing methodologies
Experience with cost optimization strategies in cloud and Kubernetes environments
Security certifications (AWS Security Specialty, CISSP, CKS, etc.)
Experience with compliance frameworks (SOC 2, ISO 27001, etc.)
Contributions to open-source Kubernetes projects or active participation in the Kubernetes community
What We Offer
Competitive salary and equity compensation
Comprehensive health, dental, and vision insurance
Flexible work arrangements
Professional development opportunities and certification support
Collaborative and inclusive team culture
Our Commitment
We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.