1

Site Reliability Engineer Jobs in Grayson, GA (NOW HIRING)

Senior Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These ...

New

Senior Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These ...

New

As a Senior Site Reliability Engineer, you'll play a pivotal role in ensuring the reliability, scalability, and performance of our infrastructure as we continue to scale and expand our operations.

New

GCP SRE

Alpharetta, GA · On-site

$55.75 - $74/hr

This role is for a Site Reliability Engineer (SRE) with a strong emphasis on Google Cloud Platform (GCP) and AI skills. The position requires robust incident management capabilities within the GCP ...

Showing results 21-40

Site Reliability Engineer information

See Grayson, GA salary details

$10

$59

$85

How much do site reliability engineer jobs pay per hour?

As of Aug 20, 2026, the average hourly pay for site reliability engineer in Grayson, GA is $59.15, according to ZipRecruiter salary data. Most workers in this role earn between $50.87 and $67.60 per hour, depending on experience, location, and employer.

What is a site reliability engineer?

A site reliability engineer specializes in site reliability engineering, or SRE, a specific branch of operations first pioneered by Google. You are responsible for ensuring that when a website decides to scale a particular feature for various users to access, it does not break the underlying software or website functions. This means you need to use analytical problem-solving skills to determine how to make specific features on a new software release work on top of existing source code.

What is a site reliability engineer?

A Site Reliability Engineer (SRE) is a professional who applies software engineering principles to infrastructure and operations problems. Their primary goal is to create scalable and highly reliable software systems, often bridging the gap between development and IT operations. SREs automate tasks, monitor system health, respond to incidents, and work to improve system reliability and performance. They also help define service level objectives (SLOs) and ensure systems meet customer expectations for uptime and availability.

What are the key skills and qualifications needed to thrive as a site reliability engineer?

To thrive as a Site Reliability Engineer, you need a strong background in computer science, systems administration, and software engineering, often supported by a degree in a technical field. Familiarity with cloud platforms (like AWS or GCP), container orchestration (such as Kubernetes), infrastructure as code (Terraform or Ansible), and monitoring tools (Prometheus, Grafana) is typically expected. Strong problem-solving skills, effective communication, and a proactive mindset help SREs excel at incident management and cross-functional collaboration. These skills are crucial for maintaining system reliability, minimizing downtime, and driving continuous improvement in complex technical environments.

What are some of the most common challenges site reliability engineers face when balancing system reliability with rapid software delivery?

Site Reliability Engineers (SREs) often navigate the challenge of maintaining highly reliable systems while supporting fast-paced software releases. This involves managing incidents, automating processes to reduce manual toil, and working closely with development teams to embed reliability into the software development lifecycle. SREs must carefully prioritize their efforts between proactive improvements and urgent, reactive fire-fighting. Effective communication and collaboration with both operations and development teams are crucial to ensuring service uptime without slowing down innovation.

What is the difference between Site Reliability Engineer vs DevOps Engineer?

AspectSite Reliability EngineerDevOps Engineer
CredentialsTypically requires a computer science degree, certifications like AWS, Google Cloud, or KubernetesSimilar credentials, often with cloud certifications and scripting skills
Work EnvironmentFocuses on maintaining and improving system reliability, often in large-scale production environmentsWorks on automation, CI/CD pipelines, and deployment processes across development and operations teams
Industry UsageCommon in tech, cloud services, and large-scale enterprise companiesWidely used in software development, cloud, and IT organizations

Both roles require strong technical skills and cloud knowledge, but SREs focus more on system reliability and uptime, while DevOps engineers emphasize automation and deployment processes. They often collaborate but have distinct primary responsibilities.

Is a site reliability engineer a stressful job?

A site reliability engineer (SRE) role can be stressful due to the responsibility of maintaining system uptime, handling incidents, and ensuring reliability under tight deadlines. The job often requires strong problem-solving skills, familiarity with monitoring tools, and the ability to work in high-pressure situations, but it also offers opportunities for skill development and process improvements.

What cities near Grayson, GA are hiring for Site Reliability Engineer jobs?

Cities near Grayson, GA with the most Site Reliability Engineer job openings:

Infographic showing various Site Reliability Engineer job openings in Grayson, GA as of August 2026, with employment types broken down into 1% As Needed, 77% Full Time, 17% Part Time, 1% Temporary, and 4% Contract. Highlights an 93% Physical, 2% Hybrid, and 5% Remote job distribution, with an average salary of $123,029 per year, or $59.1 per hour.

Senior Site Reliability Engineer

Inspire Brands

Atlanta, GA • On-site

$54.75 - $72.75/hr

Full-time

Posted 3 days ago

New


Inspire Brands rating

5.8

Company rating: 5.8 out of 10

Based on 56 frontline employees who took The Breakroom Quiz

27th of 106 rated fast food restaurants


Job description

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.

The ideal candidate has hands-on experience applying and implementing SRE principles - not just supporting production systems, but engineering reliability into them.

RESPONSIBILITIES

Reliability Engineering

  • Define and manage SLIs, SLOs, and Error Budgets for critical services
  • Drive production readiness reviews and reliability requirements into architecture and design
  • Perform capacity planning, failure mode analysis, and dependency risk assessments
  • Identify systemic reliability risks and drive remediation before they cause customer impact

Observability

  • Design monitoring, alerting, logging, and tracing solutions using modern observability tooling
  • Improve signal-to-noise ratio and reduce alert fatigue
  • Build dashboards and telemetry that reflect true service health, not just infrastructure metrics

Incident Management

  • Lead technical response for high-severity incidents
  • Drive blameless postmortems and root cause analysis focused on systemic fixes
  • Continuously improve detection, response, and recovery processes
  • Participate in an on-call rotation

Automation & Toil Reduction

  • Identify and eliminate manual, repetitive operational work through automation
  • Build self-healing systems, tooling, and scripts to reduce human intervention
  • Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green)
  • Support Infrastructure as Code (Terraform, Bicep, or similar)

Performance & Scalability

  • Conduct load testing, performance benchmarking, and bottleneck analysis
  • Partner with engineering to design systems for horizontal scalability and fault tolerance

Collaboration & Culture

  • Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)
  • Mentor engineers on SRE best practices
  • Promote a culture of engineering-driven reliability over reactive operations

EDUCATION AND EXPERIENCE QUALIFICATIONS

Required Qualifications

  • 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
  • 2+ years experience with Kubernetes and containerized workloads
  • 4-year degree in Computer Scienceor related field

Preferred Qualifications

  • Experience with chaos engineering or resiliency testing
  • Experience with high-volume, high-availability transactional systems
  • Experience with AI-assisted observability or operational automation
  • Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms

REQUIRED KNOWLEDGE, SKILLS, OR ABILITIES

  • Strong programming/scripting skills (Python, Go, Java, or Node.js)
  • Demonstrated experience defining and operating against SLOs/Error Budgets
  • Strong skills in leading incident response and root cause analysis for production systems
  • Solid understanding of distributed systems and microservices architecture
  • Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)
  • Expertise with observability platforms and monitoring strategy

This position is based in our Atlanta Support Center, with an expected on-site presence of 80%.


Inspire is a multi-brand restaurant company whose portfolio includes more than 33,300 Arby's, Baskin-Robbins, Buffalo Wild Wings, Dunkin', Jimmy John's, and SONIC restaurants worldwide. We're made up of some of the world's most iconic restaurant brands, but we're much more than just a restaurant company. We're a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple-it's an experience. At Inspire, that's our purpose: to ignite and nourish flavorful experiences.

What Inspire Brands employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Inspire Brands logo

About Inspire Brands

Sourced by ZipRecruiter

Inspire Brands Inc., located in Atlanta, GA, United States, operates in the foodservice industry as a multi-brand restaurant company, making it among the biggest restaurant companies globally. Their portfolio includes well-known restaurant brands such as Arby's, Buffalo Wild Wings, Sonic, and Jimmy John's, reflecting their commitment to innovation and quality. Founded in 2018 as a result of a consolidation of various restaurant brands under one corporate umbrella, Inspire Brands was formed with a vision to invigorate excellent brands and supercharge their long-term growth.

Industry

Food services and drinking places

Company size

10,000+ Employees

Headquarters location

Atlanta, GA, US

Year founded

2018