1

Site Reliability Engineer Manager Jobs in Georgia

Senior Site Reliability Engineer

Atlanta, GA

$54.75 - $72.75/hr

About Your Role As a Senior Site Reliability Engineer, you will work with our Infrastructure and ... Design, implement, and manage scalable systems that ensure high availability, fault tolerance, and ...

Senior Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Who We Are QGenda is redefining healthcare workforce management everywhere care is delivered. We're ... About Your Role As a Senior Site Reliability Engineer, you will work with our Infrastructure and ...

Site Reliability Engineer - Networking

Atlanta, GA

$54.75 - $72.75/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and ... Design, implementation and management of an overlay network to support 1000's of containers.

Senior Site Reliability Engineer (SRE)

Atlanta, GA · On-site

$120K - $175K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

As a Senior Site Reliability Engineer, you'll play a pivotal role in ensuring the reliability, scalability, and performance of our infrastructure as we continue to scale and expand our operations.

Senior Site Reliability Engineer (SRE)

Atlanta, GA · On-site +1

$120K - $175K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

As a Senior Site Reliability Engineer, you'll play a pivotal role in ensuring the reliability, scalability, and performance of our infrastructure as we continue to scale and expand our operations.

Site Reliability Engineer - Networking

Alpharetta, GA

$55.75 - $74/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and ... Design, implementation and management of an overlay network to support 1000's of containers.

Define and manage SLOs, SLIs, and error budgets * Build and improve CI/CD pipelines and operational ... Mentor other engineers and help set SRE standards and best practices Required Qualifications * 5+ ...

Sr. Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

Job Title : Sr. Site Reliability Engineer Location : Atlanta, GA - Hybrid Duration : 6+ Months ... Extensive/Strong AWS experience: experience in designing, deploying managing scalable/reliable ...

Define and manage SLOs, SLIs, and error budgets * Build and improve CI/CD pipelines and operational ... Mentor other engineers and help set SRE standards and best practices Required Qualifications * 5+ ...

Senior Site Reliability Engineer II

Buford, GA · On-site +1

$125K - $209K/yr

Define and manage SLOs, SLIs, and error budgets * Build and improve CI/CD pipelines and operational ... Mentor other engineers and help set SRE standards and best practices Required Qualifications * 5+ ...

Site Reliability Engineer (AWS)

Atlanta, GA · Hybrid

$54.75 - $72.75/hr

This position is under our CTO org to support SRE functions for innovation and growth for the ... Manage deployment pipelines and configuration management for consistent and reliable app ...

Showing results 41-60

Site Reliability Engineer Manager information

See Georgia salary details

$9

$53

$77

How much do site reliability engineer manager jobs pay per hour?

As of Aug 16, 2026, the average hourly pay for site reliability engineer manager in Georgia is $53.82, according to ZipRecruiter salary data. Most workers in this role earn between $46.30 and $61.49 per hour, depending on experience, location, and employer.

What is a site reliability engineer manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How does a site reliability engineer manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What are the key skills and qualifications needed to thrive as a site reliability engineer manager?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.

What are the most commonly searched types of Site Reliability Engineer jobs in Georgia?

The most popular types of Site Reliability Engineer jobs in Georgia are:

What cities in Georgia are hiring for Site Reliability Engineer Manager jobs?

Cities in Georgia with the most Site Reliability Engineer Manager job openings:

Infographic showing various Site Reliability Engineer Manager job openings in Georgia as of August 2026, with employment types broken down into 1% As Needed, 82% Full Time, 14% Part Time, 1% Temporary, and 2% Contract. Highlights an 93% Physical, 3% Hybrid, and 4% Remote job distribution, with an average salary of $111,951 per year, or $53.8 per hour.

Site Reliability Engineer (SRE) - AI Platform & Cloud

Morgan Stanley

Alpharetta, GA • On-site

$55.75 - $74/hr

Other

Re-posted 26 days ago


Morgan Stanley rating

8.4

Company rating: 8.4 out of 10

Based on 155 frontline employees who took The Breakroom Quiz

30th of 150 rated financial services


Job description

In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities.
This is a Software Engineering position at Director level, which is part of the job family responsible for developing and maintaining software solutions that support business needs.
Since 1935, Morgan Stanley is known as a global leader in financial services, always evolving and innovating to better serve our clients and our communities in more than 40 countries around the world.
Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firm's Technology principles and drives efficiency and consistency, controls, security and strong governance and promotes innovation, enabling teams to build applications that leverage AI capabilities and accelerate the adoption of AI across our businesses.
This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and harden the infrastructure that powers our AI/ML systems. You will collaborate closely with infrastructure engineering, cloud engineering, data engineering, and security teams to ensure availability, reliability, performance, and security of production AI workloads (training, inference, data pipelines) in a regulated, high-stakes financial environment.
As an SRE on the AI platform, you will bring deep operations, automation, and systems engineering skills to enable our models and pipelines to run reliably at scale, while balancing cost, security, and compliance constraints.
The ideal candidate will have strong hands-on experience supporting software platforms on any combination of the following platforms - Kubernetes, Cloud (AWS, Azure, and/or Google), API based development, REST framework, data engineering, and large-scale API Gateway environments etc. Knowledge of AIML and hands-on experience implementing solutions using Generative AI are also preferable. The candidate will have great communication skills, a team-based mentality and a strong passion for using AI to increase productivity as well as help generate new ideas for product & technical improvements.
What you'll do in the role:
  • Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)
  • Design and build automation for core platform capabilities, reducing manual toil
  • Develop and maintain infrastructure-as-code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
  • Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards
  • Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation
  • Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting
  • Optimize cost vs. performance tradeoffs in large-scale compute environments
  • Harden systems for security, compliance, auditability, and data governance
  • Collaborate across teams (cloud engineers, data engineers, infrastructure, security) to ensure safe deployment, rollout, rollback, and integration of new systems
  • Define disaster recovery (DR) strategies, backup/restore practices, fault tolerance mechanisms
  • Maintain runbooks, operational playbooks, documentation, and training materials
  • Participate in on-call rotations and respond to production incidents 24/7 as needed
  • Continuously evaluate and integrate new tools, frameworks, or technologies to enhance platform reliability

What you'll bring to the role:
  • Bachelor's or Master's degree in Computer Science or related field, or equivalent job experience
  • 5 years of production experience in SRE / Infrastructure / ops for large-scale systems
  • Strong programming/scripting skills (Python, Go, Java, or equivalent)
  • Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
  • Infrastructure-as-code (Terraform, Helm, CloudFormation, Ansible, etc.)
  • Familiarity with GPU / AI compute clusters, high-performance data storage, and distributed architectures
  • Experience with monitoring / observability / logging / alerting tools (Prometheus, Grafana, ELK / EFK, Datadog, etc.)
  • Networking & systems engineering knowledge (TCP/IP, DNS, routing, load balancing, distributed storage)
  • Solid experience in capacity planning, performance tuning, scaling, and incident response
  • Demonstrated ability to lead RCAs, deploy fixes, and drive reliability improvements
  • Experience in regulated environments (financial services, compliance, audit, security) is a strong plus
  • Excellent communication, documentation, and cross-team collaboration skills
  • Proven track record of reducing operational toil via automation

Nice to have
  • Understanding of SRE techniques.
  • Proficiency with Open Telemetry tools including Grafana, Loki, Prometheus, and Cortex.
  • Good knowledge of Microservice based architecture, industry standards, for both public and private cloud.
  • Knowledge of data pipeline technologies (Kafka, Spark, Flink, etc.)
  • Good knowledge of various DB engines (SQL, Redis, Kafka, Snowflake, etc) for cloud app storage.
  • Experience working with Generative AI development, embeddings, fine tuning of Generative AI models.
  • Experience in high-performance computing (HPC), distributed GPU cluster scheduling (e.g. Slurm, Kubernetes GPU scheduling)
  • Understanding of ModelOps/ ML Ops/ LLM Op.
  • Experience with chaos engineering, canary deployments, blue/green rollouts

We have a track record of innovation and passion for unlocking new opportunities, we help our clients raise, manage and allocate capital. We do this by offering a wide range of investment banking, securities, wealth management and asset management services.
All that we do at Morgan Stanley is driven by our five core values: do the right thing, put clients first, lead with exceptional ideas, commit to diversity and inclusion, and give back. These aren't just beliefs, they guide the decisions we make every day, ensuring we do what's best for our clients, communities and more than 80,000 employees around the world. And at the core of our success are the people who drive it - relentless collaborators and creative thinkers who are fueled by diverse thinking and experiences.
Wherever you are in our 1,200 global offices, you'll have the opportunity to work alongside the best and the brightest in an environment where you are empowered to achieve your full potential. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry.
At Morgan Stanley Alpharetta, we support the Firm's global business and functions from Wealth Management and Institutional Securities to Technology and Operations, Finance and Human Resources. With the 2020 acquisition of E-TRADE, Morgan Stanley Alpharetta grew significantly and has grown its role in our Wealth Management business helping deliver a premiere experience for the digitally inclined investor and trader. Learn more about our work and culture in Morgan Stanley Alpharetta.
Morgan Stanley's goal is to build and maintain a workforce that is diverse in experience and background but uniform in reflecting our standards of integrity and excellence. Consequently, our recruiting efforts reflect our desire to attract and retain the best and brightest from all talent pools. We want to be the first choice for prospective employees.
It is the policy of the Firm to ensure equal employment opportunity without discrimination or harassment on the basis of race, color, religion, creed, age, sex, sex stereotype, gender, gender identity or expression, transgender, sexual orientation, national origin, citizenship, disability, marital and civil partnership/union status, pregnancy, veteran or military service status, genetic information, or any other characteristic protected by law.
Morgan Stanley is an equal opportunity employer committed to diversifying its workforce (M/F/Disability/Vet).
WHAT YOU CAN EXPECT FROM MORGAN STANLEY:
At Morgan Stanley, we raise, manage and allocate capital for our clients - helping them reach their goals. We do it in a way that's differentiated - and we've done that for 90 years. Our values - putting clients first, doing the right thing, leading with exceptional ideas, committing to diversity and inclusion, and giving back - aren't just beliefs, they guide the decisions we make every day to do what's best for our clients, communities and more than 80,000 employees in 1,200 offices across 42 countries. At Morgan Stanley, you'll find an opportunity to work alongside the best and the brightest, in an environment where you are supported and empowered. Our teams are relentless collaborators and creative thinkers, fueled by their diverse backgrounds and experiences. We are proud to support our employees and their families at every point along their work-life journey, offering some of the most attractive and comprehensive employee benefits and perks in the industry. There's also ample opportunity to move about the business for those who show passion and grit in their work.
To learn more about our offices across the globe, please copy and paste into your browser.
Morgan Stanley is an equal opportunity employer committed to building and maintaining a workforce that is diverse in experience and background. Our recruiting efforts reflect our strong commitment to a culture of inclusion, where individuals are hired, developed, and advanced based on their skills and talents.
Our workforce reflects a broad cross-section of the global communities in which we operate, bringing a variety of backgrounds, talents, perspectives, and experiences.
For more information, please visit:

What Morgan Stanley employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom