1

Home Based Site Reliability Engineer Jobs in Georgia

Site Reliability Engineer (SRE)

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

The ideal candidate has strong experience supporting Java-based applications running on Amazon EKS ... SRE, DevOps, Production Support, or Site Reliability Engineering role. * Strong experience ...

Site Reliability Engineer

Alpharetta, GA ยท On-site

$55.75 - $74/hr

Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 ... Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to ...

Site Reliability Engineer - SRE

Atlanta, GA ยท On-site

$54.25 - $72/hr

Role: Site Reliability Engineer * Location: Atlanta, GA OR Dallas OR Austin, TX * Duration ... Proficient in a Linux or Unix based environment. * Proficiency in supporting a 24x7 operation.

Site Reliability Engineer (SRE)

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Job Summary : eTeam is a company seeking a Site Reliability Engineer (SRE) for a contract position ... Responsibilities : โ€ข Develop Ansible, Python based automation โ€ข Automate provisioning, patching ...

Site Reliability Engineer - SRE

Atlanta, GA ยท On-site +1

$54.25 - $72/hr

Role: Site Reliability Engineer * Location: Atlanta, GA OR Dallas OR Austin, TX * Duration ... Proficient in a Linux or Unix based environment. * Proficiency in supporting a 24x7 operation.

Site Reliability Engineer

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 ... Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to ...

Site Reliability Engineer

Alpharetta, GA ยท On-site

$55.75 - $74/hr

Atlanta-based Incident IQ is the leading workflow management platform built exclusively for K-12 ... Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to ...

Site Reliability Engineer

Norcross, GA ยท Hybrid

$53.50 - $71/hr

Build and operate Azure-based platforms (AKS, networking, security, monitoring) * Implement GitOps ... SRE, DevOps, or Platform Engineering * Strong Azure expertise (AKS, monitoring, networking ...

New

Site Reliability Engineer

Atlanta, GA ยท On-site +1

$100K - $120K/yr

Overview The Site Reliability Engineer is a key force behind improving Origami's time to resolution ... Net based web applications to identify bugs/performance challenges. * Solid knowledge of SaaS ...

Site Reliability Engineer

Norcross, GA ยท On-site

$53.50 - $71/hr

Build and operate Azure-based platforms (AKS, networking, security, monitoring) * Implement GitOps ... SRE, DevOps, or Platform Engineering * Strong Azure expertise (AKS, monitoring, networking ...

Site Reliability Engineer

Atlanta, GA ยท On-site

$100 - $120/hr

Overview The Site Reliability Engineer is a key force behind improving Origami's time to resolution ... Net based web applications to identify bugs/performance challenges. * Solid knowledge of SaaS ...

Site Reliability Engineer

Alpharetta, GA ยท On-site

$55.75 - $74/hr

I have an opportunity for a " Site Reliability Engineer " - Alpharetta, GA (Onsite). and I am looking for a candidate who can join Immediately if you are interested, reply to me with your updated ...

next page

Showing results 1-20

Home Based Site Reliability Engineer information

What is a home based site reliability engineer?

A Home Based Site Reliability Engineer is an IT professional who works remotely to ensure the reliability, availability, and performance of software systems and infrastructure. They blend software engineering and systems administration skills to automate processes, monitor system health, and respond to incidents. Working from home, they use collaboration tools to communicate with teams, manage cloud resources, and implement best practices for uptime and scalability. Their role is crucial to maintaining seamless digital experiences and minimizing downtime.

What are the key skills and qualifications needed to thrive as a home based site reliability engineer?

To thrive as a Home Based Site Reliability Engineer, you need a strong background in computer science, experience with systems administration, and proficiency in programming languages like Python or Go, often backed by a relevant degree or certifications. Familiarity with cloud platforms (AWS, GCP, or Azure), CI/CD pipelines, containerization tools (Docker, Kubernetes), and monitoring systems is typically required. Strong problem-solving abilities, effective communication, and self-motivation are vital soft skills for collaborating remotely and responding to incidents efficiently. These competencies ensure high system reliability, rapid issue resolution, and smooth collaboration in distributed work environments.

How does being a home based site reliability engineer impact collaboration with on-site teams and incident response?

As a home-based Site Reliability Engineer, you will frequently rely on digital communication tools to collaborate with both remote and on-site teams. Effective virtual coordination is essential, especially during incident response, where rapid, clear communication can significantly impact resolution times. Many organizations use platforms like Slack, PagerDuty, or Zoom to facilitate real-time collaboration, ensuring that remote SREs remain integrated in monitoring, troubleshooting, and post-incident reviews. While remote work offers flexibility, it also requires strong self-motivation and proactive communication to stay aligned with team goals and maintain high system reliability.

What is the difference between Home Based Site Reliability Engineer vs Cloud Operations Engineer?

AspectHome Based Site Reliability EngineerCloud Operations Engineer
CredentialsCertifications like AWS, Google Cloud, or Azure; SRE-specific trainingCloud platform certifications; DevOps and infrastructure knowledge
Work EnvironmentRemote, often with distributed teamsRemote or on-premises, depending on company
Industry UsageTech, SaaS, and cloud service providersCloud service providers, enterprise IT

Home Based Site Reliability Engineers focus on maintaining system reliability, scalability, and automation primarily in cloud environments, often working remotely. Cloud Operations Engineers also manage cloud infrastructure but may have a broader focus on deployment, monitoring, and incident response. Both roles require cloud certifications and share similar work environments, but SREs emphasize reliability engineering principles.

Are home based site reliability engineers in demand?

Home-based site reliability engineers are in high demand due to the increasing reliance on cloud services and digital infrastructure. They are sought after for their skills in monitoring, automation, and using tools like Kubernetes and cloud platforms, with remote work options becoming more common in the industry.

How much do home based site reliability engineers get paid?

Home-based site reliability engineers typically earn between $80,000 and $150,000 annually, depending on experience, location, and company size. Salaries often increase with proficiency in cloud platforms, automation tools, and monitoring systems, and may include benefits like flexible schedules and remote work allowances.

What are the most commonly searched types of Site Reliability Engineer jobs in Georgia?

The most popular types of Site Reliability Engineer jobs in Georgia are:

What job categories do people searching Home Based Site Reliability Engineer jobs in Georgia look for?

The top searched job categories for Home Based Site Reliability Engineer jobs in Georgia are:

What cities in Georgia are hiring for Home Based Site Reliability Engineer jobs?

Cities in Georgia with the most Home Based Site Reliability Engineer job openings:

Infographic showing various Home Based Site Reliability Engineer job openings in Georgia as of June 2026, with employment types broken down into 52% Full Time, 46% Part Time, 1% Temporary, and 1% Contract. Highlights an 94% Physical, 3% Hybrid, and 3% Remote job distribution.

Site Reliability Engineer (SRE)

Atlanta, GA โ€ข On-site

ATLANTICUS
Finance and Insuranceย โ€ขย 201 - 500 employees

$54.75 - $72.75/hr

Full-time

Medical, Retirement, PTO

Posted 17 days ago


Job description

Site Reliability Engineer (SRE)

When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage entrepreneurial thinking to empower our customers toward financial well-being.ย 

Atlanticusโ„ข technology enables bank, retail, and healthcare partners to offer more inclusive financial services to everyday Americansย through the use ofย proprietary analytics. We apply the experience gained and infrastructure built from servicing overย 20 million customersย and overย $40 billionย in consumer loans over more than 25 years of operating history to support lenders thatย originateย a range of consumer loan products. These products include retail and healthcare, private label credit and general-purpose credit cards marketed through our omnichannel platform, including retail point-of-sale, healthcare point-of-care, direct mail solicitation, digital marketing, and partnerships with third parties. Additionally, through our Auto Finance subsidiary, Atlanticus serves the individual needs of automotive dealers and automotive non-prime financial organizations with multiple financing and service programs.ย 

Office Locations available for this role includeย 

  • Austin, TXย โ€“ Situated in The Domain, a vibrant tech hub with park-like surroundings, top restaurants, and convenient parking, perfect for post-work socializing.ย 
  • Atlanta, GAย โ€“ Located in the Queen Building (King & Queen Towers, Sandy Springs), with easy access to I-285, GA-400, and a free shuttle to MARTA.ย 

Work Cultureย 

We foster a collaborative, innovative environment where everyone contributes to building something meaningful.ย Youโ€™llย be empowered to lead, grow, and make an impact.ย 

The Roleย 

We are seeking a Site Reliability Engineer (SRE)ย to ensure the reliability, availability, performance, and operational excellence of our cloud-native applications running on AWS. This is a hands-on role responsible for monitoring and supporting production systems, automating operational tasks, managing deployments, and driving continuous improvements in system stability.

The ideal candidate has strong experience supporting Java-based applications running on Amazon EKS, a solid understanding of AWS infrastructure, and expertise with observability platforms such as Datadogย and Splunk. This role requires participation in a 24x7 production support and on-call rotation, working closely with Development, DevOps, IT Ops, Database, Network, and Security teamsย to maintain highly available production services.

The successful candidate should be passionate about automation, troubleshooting complex production issues, improving application reliability, and leveraging AI-powered tools to enhance operational efficiency.

Key Responsibilities

  • Provide 24x7 production supportย through an on-call rotation to ensure application availability and rapid incident response.
  • Continuously monitor production applications, infrastructure, and platform health using Datadog, Splunk, CloudWatch, and other monitoring tools.
  • Respond to production incidents, troubleshoot issues, and restore services while minimizing customer impact.
  • Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
  • Deploy and support Java-based applications running on Docker and Amazon EKS using CI/CD pipelines.
  • Execute production deployments, application releases, hotfixes, and rollbacks following change management processes.
  • Monitor and manage scheduled application jobs, batch processes, and integrations to ensure successful execution.
  • Troubleshoot Java application issues using logs, JVM metrics, thread dumps, heap dumps, and application performance metrics.
  • Analyze application, infrastructure, and Kubernetes logs using Splunk and Datadog to identify performance bottlenecks and operational issues.
  • Develop automation scripts using Python, Bash, or similar scripting languages to eliminate repetitive operational tasks.
  • Build self-healing and automated operational processes to improve system reliability and reduce manual intervention.
  • Support Kubernetes (Amazon EKS) environments, including troubleshooting pods, deployments, networking, ingress, and scaling issues.
  • Maintain and improve dashboards, alerts, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational runbooks.
  • Partner with Development teams to improve application reliability, resiliency, scalability, and performance.
  • Continuously improve operational processes, monitoring coverage, automation, and deployment practices.
  • Leverage AI-powered engineering tools and agentic AI capabilities to improve monitoring, incident response, automation, and operational efficiency.

Youโ€™reย a great fit if you haveย 

  • 5+ years of experience supporting production applications in an SRE, DevOps, Production Support, or Site Reliability Engineering role.
  • Strong experience supporting Java-based applications in production environments.
  • Hands-on experience with AWS services including EKS, EC2, ALB/NLB, RDS, IAM, Route 53, CloudWatch, S3, and VPC.
  • Experience with Kubernetes (Amazon EKS), Docker, and containerized application deployments.
  • Strong experience using Datadog / Splunk for infrastructure monitoring, APM, troubleshooting, dashboards, alerting, and log analysis.
  • Experience performing production deployments through CI/CD pipelines (Jenkins, GitHub Actions, Argo CD, or similar).
  • Experience supporting MySQL and Oracle databases from an application support perspective.
  • Proficiency in Python, Bash, or other scripting languages for automation.
  • Strong Linux system administration and troubleshooting skills.
  • Excellent troubleshooting skills across distributed applications, networking, and cloud infrastructure.
  • Knowledge of networking fundamentals including DNS, TCP/IP, HTTP/HTTPS, TLS, load balancing, and firewalls.
  • Experience with incident management, problem management, and change management processes.
  • Experience using AI-assisted development tools or agentic AI systems to improve operational efficiency.

Preferredย ย 

  • Experience with Helm and GitOps deployment models.
  • Experience with Terraform or Infrastructure as Code.
  • Familiarity with Prometheus, Grafana, or OpenTelemetry.
  • Experience supporting microservices architectures.
  • Knowledge of JVM tuning and Java performance optimization.
  • Experience with AWS Auto Scaling, Karpenter, or Cluster Autoscaler.

Whyย Youโ€™llย Love Working Hereย 

Thisย isnโ€™tย just aย job,ย itโ€™sย a place to lead, grow, and thrive. If you believe in your skills and drive,ย weโ€™llย provideย the resourcesย and support to help you succeed.ย 

Benefits includeย 

  • Generous PTO and holiday scheduleย 
  • 401(k) with company matchย 
  • Employee stock purchase planย 
  • Ongoing training (lunch & learns, financial and health webinars)ย 
  • Team volunteer outingsย 

Atlanticus is an equal opportunity employer. All qualified applicants will receive consideration without regard to race, religion, gender, sexual orientation, age, veteran status, disability, or other protected status.ย