1

Site Reliability Engineer Jobs in Grayson, GA (NOW HIRING)

Site Reliability Engineer (SRE)

Atlanta, GA · On-site

$54.75 - $72.75/hr

  • Medical

  • Retirement

  • PTO

Site Reliability Engineer (SRE) When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class ...

Site Reliability Engineer

Alpharetta, GA · On-site

$55.75 - $74/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first ...

Site Reliability Engineer

Alpharetta, GA

$55.75 - $74/hr

Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first ...

Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first ...

Site Reliability engineer (SRE)

Atlanta, GA · On-site

$54.75 - $72.75/hr

Site Reliability engineer(SRE) Location: Atlanta, GA ( Hybrid - 3days Office - 2 days WFH) Duration: C2H : Dynatrace App dynamics ACI (Advanced Computing International) is a Global Technology ...

Site Reliability Engineer

Alpharetta, GA · On-site

$55.75 - $74/hr

I have an opportunity for a " Site Reliability Engineer " - Alpharetta, GA (Onsite). and I am looking for a candidate who can join Immediately if you are interested, reply to me with your updated ...

Site Reliability Engineer

Atlanta, GA · On-site +1

$100K - $120K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Overview The Site Reliability Engineer is a key force behind improving Origami's time to resolution and advancing overall site reliability and scalability. This person participates in efforts to ...

Site Reliability Engineer

Atlanta, GA · On-site

$100 - $120/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Overview The Site Reliability Engineer is a key force behind improving Origami's time to resolution and advancing overall site reliability and scalability. This person participates in efforts to ...

Site Reliability Engineer

Alpharetta, GA · On-site

$65 - $75/hr

  • Medical

  • Dental

  • Vision

Title: Senior Site Reliability Engineer Location: Alpharetta, GA Duration: 6-12+ Months About the Role We're seeking an experienced Senior Site Reliability Engineer to join our team and play a ...

Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

This SRE will be supporting a full stack React and Java application that also has salesforce integrations and a KTLO team. They are looking for someone strong in Google Cloud Platform and able to ...

SRE Lead/ Architect

Atlanta, GA · On-site

$54.75 - $72.75/hr

Job Title: SRE Lead/Architect Location: Atlanta, GA - Hybrid (Thur to next wed (Alternate weeks)) Contract Role Role Summary: Mandatory skills are Observability, Resiliency, Chaos engineering, strong ...

next page

Showing results 1-20

Site Reliability Engineer information

See Grayson, GA salary details

$10

$59

$85

How much do site reliability engineer jobs pay per hour?

As of Aug 19, 2026, the average hourly pay for site reliability engineer in Grayson, GA is $59.15, according to ZipRecruiter salary data. Most workers in this role earn between $50.87 and $67.60 per hour, depending on experience, location, and employer.

What is a site reliability engineer?

A site reliability engineer specializes in site reliability engineering, or SRE, a specific branch of operations first pioneered by Google. You are responsible for ensuring that when a website decides to scale a particular feature for various users to access, it does not break the underlying software or website functions. This means you need to use analytical problem-solving skills to determine how to make specific features on a new software release work on top of existing source code.

What is a site reliability engineer?

A Site Reliability Engineer (SRE) is a professional who applies software engineering principles to infrastructure and operations problems. Their primary goal is to create scalable and highly reliable software systems, often bridging the gap between development and IT operations. SREs automate tasks, monitor system health, respond to incidents, and work to improve system reliability and performance. They also help define service level objectives (SLOs) and ensure systems meet customer expectations for uptime and availability.

What are the key skills and qualifications needed to thrive as a site reliability engineer?

To thrive as a Site Reliability Engineer, you need a strong background in computer science, systems administration, and software engineering, often supported by a degree in a technical field. Familiarity with cloud platforms (like AWS or GCP), container orchestration (such as Kubernetes), infrastructure as code (Terraform or Ansible), and monitoring tools (Prometheus, Grafana) is typically expected. Strong problem-solving skills, effective communication, and a proactive mindset help SREs excel at incident management and cross-functional collaboration. These skills are crucial for maintaining system reliability, minimizing downtime, and driving continuous improvement in complex technical environments.

What are some of the most common challenges site reliability engineers face when balancing system reliability with rapid software delivery?

Site Reliability Engineers (SREs) often navigate the challenge of maintaining highly reliable systems while supporting fast-paced software releases. This involves managing incidents, automating processes to reduce manual toil, and working closely with development teams to embed reliability into the software development lifecycle. SREs must carefully prioritize their efforts between proactive improvements and urgent, reactive fire-fighting. Effective communication and collaboration with both operations and development teams are crucial to ensuring service uptime without slowing down innovation.

What is the difference between Site Reliability Engineer vs DevOps Engineer?

AspectSite Reliability EngineerDevOps Engineer
CredentialsTypically requires a computer science degree, certifications like AWS, Google Cloud, or KubernetesSimilar credentials, often with cloud certifications and scripting skills
Work EnvironmentFocuses on maintaining and improving system reliability, often in large-scale production environmentsWorks on automation, CI/CD pipelines, and deployment processes across development and operations teams
Industry UsageCommon in tech, cloud services, and large-scale enterprise companiesWidely used in software development, cloud, and IT organizations

Both roles require strong technical skills and cloud knowledge, but SREs focus more on system reliability and uptime, while DevOps engineers emphasize automation and deployment processes. They often collaborate but have distinct primary responsibilities.

Is a site reliability engineer a stressful job?

A site reliability engineer (SRE) role can be stressful due to the responsibility of maintaining system uptime, handling incidents, and ensuring reliability under tight deadlines. The job often requires strong problem-solving skills, familiarity with monitoring tools, and the ability to work in high-pressure situations, but it also offers opportunities for skill development and process improvements.

What cities near Grayson, GA are hiring for Site Reliability Engineer jobs?

Cities near Grayson, GA with the most Site Reliability Engineer job openings:

Infographic showing various Site Reliability Engineer job openings in Grayson, GA as of August 2026, with employment types broken down into 1% As Needed, 77% Full Time, 17% Part Time, 1% Temporary, and 4% Contract. Highlights an 93% Physical, 2% Hybrid, and 5% Remote job distribution, with an average salary of $123,029 per year, or $59.1 per hour.

Site Reliability Engineer (SRE)

ATLANTICUS

Atlanta, GA • On-site

$54.75 - $72.75/hr

Full-time

Medical, Retirement, PTO

Posted 12 days ago


Job description

Site Reliability Engineer (SRE)

When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage entrepreneurial thinking to empower our customers toward financial well-being. 

Atlanticus™ technology enables bank, retail, and healthcare partners to offer more inclusive financial services to everyday Americans through the use of proprietary analytics. We apply the experience gained and infrastructure built from servicing over 20 million customers and over $40 billion in consumer loans over more than 25 years of operating history to support lenders that originate a range of consumer loan products. These products include retail and healthcare, private label credit and general-purpose credit cards marketed through our omnichannel platform, including retail point-of-sale, healthcare point-of-care, direct mail solicitation, digital marketing, and partnerships with third parties. Additionally, through our Auto Finance subsidiary, Atlanticus serves the individual needs of automotive dealers and automotive non-prime financial organizations with multiple financing and service programs. 

Office Locations available for this role include 

  • Austin, TX – Situated in The Domain, a vibrant tech hub with park-like surroundings, top restaurants, and convenient parking, perfect for post-work socializing. 
  • Atlanta, GA – Located in the Queen Building (King & Queen Towers, Sandy Springs), with easy access to I-285, GA-400, and a free shuttle to MARTA. 

Work Culture 

We foster a collaborative, innovative environment where everyone contributes to building something meaningful. You’ll be empowered to lead, grow, and make an impact. 

The Role 

We are seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational excellence of our cloud-native applications running on AWS. This is a hands-on role responsible for monitoring and supporting production systems, automating operational tasks, managing deployments, and driving continuous improvements in system stability.

The ideal candidate has strong experience supporting Java-based applications running on Amazon EKS, a solid understanding of AWS infrastructure, and expertise with observability platforms such as Datadog and Splunk. This role requires participation in a 24x7 production support and on-call rotation, working closely with Development, DevOps, IT Ops, Database, Network, and Security teams to maintain highly available production services.

The successful candidate should be passionate about automation, troubleshooting complex production issues, improving application reliability, and leveraging AI-powered tools to enhance operational efficiency.

Key Responsibilities

  • Provide 24x7 production support through an on-call rotation to ensure application availability and rapid incident response.
  • Continuously monitor production applications, infrastructure, and platform health using Datadog, Splunk, CloudWatch, and other monitoring tools.
  • Respond to production incidents, troubleshoot issues, and restore services while minimizing customer impact.
  • Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
  • Deploy and support Java-based applications running on Docker and Amazon EKS using CI/CD pipelines.
  • Execute production deployments, application releases, hotfixes, and rollbacks following change management processes.
  • Monitor and manage scheduled application jobs, batch processes, and integrations to ensure successful execution.
  • Troubleshoot Java application issues using logs, JVM metrics, thread dumps, heap dumps, and application performance metrics.
  • Analyze application, infrastructure, and Kubernetes logs using Splunk and Datadog to identify performance bottlenecks and operational issues.
  • Develop automation scripts using Python, Bash, or similar scripting languages to eliminate repetitive operational tasks.
  • Build self-healing and automated operational processes to improve system reliability and reduce manual intervention.
  • Support Kubernetes (Amazon EKS) environments, including troubleshooting pods, deployments, networking, ingress, and scaling issues.
  • Maintain and improve dashboards, alerts, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational runbooks.
  • Partner with Development teams to improve application reliability, resiliency, scalability, and performance.
  • Continuously improve operational processes, monitoring coverage, automation, and deployment practices.
  • Leverage AI-powered engineering tools and agentic AI capabilities to improve monitoring, incident response, automation, and operational efficiency.

You’re a great fit if you have 

  • 5+ years of experience supporting production applications in an SRE, DevOps, Production Support, or Site Reliability Engineering role.
  • Strong experience supporting Java-based applications in production environments.
  • Hands-on experience with AWS services including EKS, EC2, ALB/NLB, RDS, IAM, Route 53, CloudWatch, S3, and VPC.
  • Experience with Kubernetes (Amazon EKS), Docker, and containerized application deployments.
  • Strong experience using Datadog / Splunk for infrastructure monitoring, APM, troubleshooting, dashboards, alerting, and log analysis.
  • Experience performing production deployments through CI/CD pipelines (Jenkins, GitHub Actions, Argo CD, or similar).
  • Experience supporting MySQL and Oracle databases from an application support perspective.
  • Proficiency in Python, Bash, or other scripting languages for automation.
  • Strong Linux system administration and troubleshooting skills.
  • Excellent troubleshooting skills across distributed applications, networking, and cloud infrastructure.
  • Knowledge of networking fundamentals including DNS, TCP/IP, HTTP/HTTPS, TLS, load balancing, and firewalls.
  • Experience with incident management, problem management, and change management processes.
  • Experience using AI-assisted development tools or agentic AI systems to improve operational efficiency.

Preferred  

  • Experience with Helm and GitOps deployment models.
  • Experience with Terraform or Infrastructure as Code.
  • Familiarity with Prometheus, Grafana, or OpenTelemetry.
  • Experience supporting microservices architectures.
  • Knowledge of JVM tuning and Java performance optimization.
  • Experience with AWS Auto Scaling, Karpenter, or Cluster Autoscaler.

Why You’ll Love Working Here 

This isn’t just a job, it’s a place to lead, grow, and thrive. If you believe in your skills and drive, we’ll provide the resources and support to help you succeed. 

Benefits include 

  • Generous PTO and holiday schedule 
  • 401(k) with company match 
  • Employee stock purchase plan 
  • Ongoing training (lunch & learns, financial and health webinars) 
  • Team volunteer outings 

Atlanticus is an equal opportunity employer. All qualified applicants will receive consideration without regard to race, religion, gender, sexual orientation, age, veteran status, disability, or other protected status.