1

Site Reliability Engineer Intern Jobs in Atlanta, GA

SRE

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems and practices that ensure the reliability, scalability ...

New

Site Reliability Engineer - SRE

Atlanta, GA

$54.75 - $72.75/hr

Site Reliability Engineer * Location: Atlanta, GA OR Dallas OR Austin, TX * Duration: Long Term or 6+ Months contract to Hire Note: Remote Possible, however candidates will move to work onsite/Hybrid ...

Site Reliability Engineer - SRE

Atlanta, GA ยท On-site

$54.25 - $72/hr

Site Reliability Engineer * Location: Atlanta, GA OR Dallas OR Austin, TX * Duration: Long Term or 6+ Months contract to Hire Note: Remote Possible, however candidates will move to work onsite/Hybrid ...

Site Reliability Engineer (SRE)

Atlanta, GA

$54.75 - $72.75/hr

Site Reliability Engineer (SRE) When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class ...

Site Reliability Engineer - SRE

Atlanta, GA

$54.25 - $72/hr

Site Reliability Engineer * Location: Atlanta, GA OR Dallas OR Austin, TX * Duration: Long Term or 6+ Months contract to Hire Note: Remote Possible, however candidates will move to work onsite/Hybrid ...

Site Reliability Engineer

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Pyramid Consulting, Inc. is seeking a Site Reliability Engineer for a 12+ months contract opportunity in Atlanta, GA. The role involves engineering software within an AWS cloud infrastructure and ...

Site Reliability Engineer (SRE)

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Job Summary : eTeam is a company seeking a Site Reliability Engineer (SRE) for a contract position. The role involves developing automation solutions using Ansible and Python, as well as maintaining ...

Site Reliability Engineer

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Site Reliability Engineer (SRE) Location: Atlanta, Georgia (Hybrid) Role Summary We are seeking an experienced Technology Consultant Site Reliability Engineer (SRE) with strong hands-on expertise in ...

New

Site Reliability Engineer

Alpharetta, GA

$55.75 - $74/hr

Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first ...

Site Reliability Engineer

Atlanta, GA

$54.75 - $72.75/hr

Site Reliability Engineer (SRE) Overview: We are looking for a Site Reliability Engineer (SRE) to join our Engineering team. This is a build-it-from-zero role at startup speed. You're our first ...

Site Reliability engineer (SRE)

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Site Reliability engineer(SRE) Location: Atlanta, GA ( Hybrid - 3days Office - 2 days WFH) Duration: C2H : Dynatrace App dynamics ACI (Advanced Computing International) is a Global Technology ...

SRE Engineer

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Job Title: SRE Location: Atlanta GA (Hybrid) Responsible for designing, building, and evolving the foundational systems and practices that ensure the reliability, scalability, performance, and ...

New

We are seeking a Senior Site Reliability Engineer (SRE) to build, automate, and support highly scalable cloud-native platforms and digital applications. This role is responsible for improving system ...

New

Site Reliability Engineer

Alpharetta, GA ยท On-site

$55.75 - $74/hr

I have an opportunity for a " Site Reliability Engineer " - Alpharetta, GA (Onsite). and I am looking for a candidate who can join Immediately if you are interested, reply to me with your updated ...

Site Reliability Engineer

Atlanta, GA ยท On-site

$100 - $120/hr

Overview The Site Reliability Engineer is a key force behind improving Origami's time to resolution and advancing overall site reliability and scalability. This person participates in efforts to ...

next page

Showing results 1-20

Site Reliability Engineer Intern information

See Atlanta, GA salary details

$10

$61

$88

How much do site reliability engineer intern jobs pay per hour?

As of Sep 2, 2026, the average hourly pay for site reliability engineer intern in Atlanta, GA is $61.30, according to ZipRecruiter salary data. Most workers in this role earn between $52.69 and $70.05 per hour, depending on experience, location, and employer.

What is a site reliability engineer intern?

A Site Reliability Engineer (SRE) Intern supports the reliability, scalability, and performance of software systems by assisting with automation, monitoring, and incident response. They work closely with development and operations teams to improve system resilience and efficiency. Typical tasks include writing scripts, analyzing system logs, and contributing to documentation or tooling improvements. This role helps interns gain hands-on experience in infrastructure management, cloud services, and DevOps practices.

What types of projects or tasks does a site reliability engineer intern typically work on?

As a Site Reliability Engineer Intern, you may work on projects like automating operational processes, building monitoring dashboards, troubleshooting incidents, and helping improve infrastructure reliability. Interns often collaborate closely with full-time engineers on tasks such as writing scripts to automate deployments, responding to on-call alerts, or optimizing system performance. You'll gain practical experience using industry-standard tools and practices while learning about uptime, scalability, and reliability workflows. It's a great opportunity to develop hands-on skills and understand how large-scale systems are maintained in a professional environment.

What are the key skills and qualifications needed to thrive as a site reliability engineer intern, and why are they important?

To thrive as a Site Reliability Engineer Intern, you should have a solid understanding of computer science fundamentals, programming (such as Python or Go), and basic networking concepts, often supported by relevant coursework or prior technical internships. Familiarity with cloud platforms (like AWS or GCP), containerization tools (e.g., Docker, Kubernetes), and version control systems (Git) is highly valuable. Strong problem-solving abilities, effective communication, and a collaborative mindset help interns stand out in team settings. These skills are crucial for learning quickly, contributing to the team's reliability initiatives, and ensuring highly available and scalable system operations.

What are the most commonly searched types of Site Reliability Engineer jobs in Atlanta, GA?

The most popular types of Site Reliability Engineer jobs in Atlanta, GA are:

What are popular job titles related to Site Reliability Engineer Intern jobs in Atlanta, GA?

For Site Reliability Engineer Intern jobs in Atlanta, GA, the most frequently searched job titles are:

What job categories do people searching Site Reliability Engineer Intern jobs in Atlanta, GA look for?

The top searched job categories for Site Reliability Engineer Intern jobs in Atlanta, GA are:

What cities near Atlanta, GA are hiring for Site Reliability Engineer Intern jobs?

Cities near Atlanta, GA with the most Site Reliability Engineer Intern job openings:

Infographic showing various Site Reliability Engineer Intern job openings in Atlanta, GA as of August 2026, with employment types broken down into 1% As Needed, 82% Full Time, 12% Part Time, 1% Temporary, 3% Contract, and 1% Nights. Highlights an 94% Physical, 2% Hybrid, and 4% Remote job distribution, with an average salary of $127,500 per year, or $61.3 per hour.

$54.75 - $72.75/hr

Other

Posted yesterday

New


Job description

Job Title: SRE Leader

Location: Atlanta GA

Attached the JD

Role Summary:

As an SRE Architect, you will be a pivotal technical leader responsible for designing, building, and evolving the foundational systems and practices that ensure the reliability, scalability, performance, and efficiency of our critical services. Moving beyond day-to-day operations, you will focus on the strategic architectural direction of SRE function, defining standards, blueprints, and frameworks that enable development teams and fellow SRE operations team to build and operate highly resilient systems. Leverage deep expertise in software engineering, distributed systems, cloud infrastructure, and SRE principles to influence technology choices, establish best practices, and foster a proactive culture of reliability across the organization and much beyond observability pillar.

Key Responsibilities:

    1. Strategy & Design: Architect and design highly available, scalable, secure, and cost-effective infrastructure and application patterns on AWS
    2. and evangelize SRE best practices, standards, and blueprints for service design, deployment, monitoring, and operational readiness across the engineering organization
    3. current observability implementation to identify gaps and define steps to reach next level maturity of observability setup to provide deep insights into system health and behaviour
    4. overall maturity lead the definition and implementation strategy for Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets for critical services
    1. Architecture & Automation: Design solutions to systematically reduce operational toil through automation and improved system design
    2. current SRE tools and automation frameworks (e.g., CI/CD pipelines, Infrastructure as Code modules, automated incident remediation, chaos engineering platforms) and suggest enhancement that will help overall enhancement of capability
    3. prototype, and recommend new technologies, tools, and methodologies to enhance system reliability, developer productivity, and operational efficiency
    1. Leadership & Consultation: Act as a senior technical advisor and subject matter expert on reliability, scalability, and performance for development and platform teams
    2. architectural guidance during the design phase of new services and features to ensure reliability principles are embedded early (shift-left)
    1. and coach other SREs and engineers, fostering technical excellence and adherence to SRE principles
    1. architectural reviews and production readiness assessments for critical systems
  1. Resilience:
    1. blameless postmortems for significant incidents, ensuring root causes are identified and systemic architectural improvements are prioritized and implemented
    2. and advocate for resilience patterns (e.g., circuit breaking, rate limiting, graceful degradation, chaos engineering) within applications and infrastructure.
  1. AI-Driven Operations & Intelligence
  2. Define and drive the organization's AIOps strategy, leveraging AI/ML and Generative AI capabilities to improve observability, incident management, root cause analysis, capacity forecasting, and operational efficiency.
  3. Design and implement intelligent operational platforms that use AI agents, knowledge graphs, telemetry analytics, and automation frameworks to proactively detect, diagnose, and remediate production issues.
  4. Architect AI-powered production intelligence solutions that correlate logs, metrics, traces, configuration data, deployment events, CMDB, cloud services, and dependency mappings to generate actionable operational insights.
  5. Establish architectural patterns for AI-assisted incident triage, impact analysis, service dependency intelligence, and automated remediation workflows.
  6. Lead the adoption of Agentic AI frameworks, MCP (Model Context Protocol) integrations, and enterprise AI platforms to improve developer productivity and operational effectiveness.
  7. Define governance frameworks for AI usage in production operations including explainability, auditability, security, risk management, and human-in-the-loop controls.
  8. Partner with Data Engineering, Platform Engineering, and Application teams to develop AI-enabled reliability use cases and operational copilots.
  9. Establish mechanisms to continuously evaluate and improve AI model effectiveness, hallucination mitigation, operational accuracy, and reliability of AI-driven decision making.

Required Qualifications:

  • Proven experience in an architectural role, designing solutions for reliability, scalability, and performance
  • Deep understanding and practical application of SRE principles (SLIs/SLOs, error budgets, toil reduction, automation, incident management, postmortems)
  • Experience designing and implementing AIOps, AI-powered observability, or intelligent automation solutions to improve incident detection, root cause analysis, operational efficiency, and service reliability.
  • Working knowledge of Generative AI, AI agents, MCP (Model Context Protocol), RAG architectures, and enterprise AI platforms, with experience building or operationalizing AI-enabled engineering tools in production environments.
  • Expertise in cloud computing platforms (e.g., AWS) including infrastructure, networking, and security services
  • Strong experience with containerization and orchestration technologies (Kubernetes, Docker, serverless computing)
  • Solid experience designing and implementing observability solutions (e.g., Dynatrace, Prometheus, Grafana, ELK/EFK Stack, Jaeger, OpenTelemetry)
  • Strong programming/scripting skills (e.g., Python, Go, Bash) for automation and tool development
  • Excellent analytical, problem-solving, and strategic thinking skills.
  • Strong communication, collaboration, and leadership skills with the ability to influence technical direction across teams

Preferred Qualifications:

  • Experience designing and implementing chaos engineering practices and platforms

Syed Moiz

Business Development & Sr Delivery Manager

EXATECH INC.

Global IT Consulting | AI | Software Engineering | Data & Analytics | Staff Augmentation

Email: