1

Amazon Sre Jobs (NOW HIRING)

Site Reliability Engineer (SRE)

Atlanta, GA · On-site

$54.75 - $72.75/hr

  • Medical

  • Retirement

  • PTO

Support Kubernetes (Amazon EKS) environments, including troubleshooting pods, deployments ... SRE, DevOps, Production Support, or Site Reliability Engineering role. * Strong experience ...

Site Reliability Engineer (SRE)

Washington, DC · On-site

$64.50 - $85.75/hr

Define and monitor SRE metrics including SLIs, SLOs, and error budgets. * Perform performance ... Strong understanding of container technologies including Docker, Kubernetes, and Amazon ECS.

Site Reliability Engineer (SRE)

Chelsea, VT

$57.50 - $76.50/hr

Define and monitor SRE metrics including SLIs, SLOs, and error budgets. * Perform performance ... Strong understanding of container technologies including Docker, Kubernetes, and Amazon ECS.

Site Reliability Engineer

Frederick, MD · On-site

$56.75 - $75.25/hr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

... Amazon AWS, Google GCP and Microsoft Azure * Must have 4+ years with CI/CD and automation tools ... and SRE use cases * Must have proficiency to debug or troubleshoot and/or deploying SQL and/or ...

Site Reliability Engineer

Frederick, MD · Hybrid

$56.75 - $75.25/hr

... Amazon AWS, Google GCP and Microsoft Azure * Must have 4+ years with CI/CD and automation tools ... and SRE use cases * Must have proficiency to debug or troubleshoot and/or deploying SQL and/or ...

Site Reliability Engineer

Frederick, MD · Hybrid

$56.75 - $75.25/hr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

... Amazon AWS, Google GCP and Microsoft Azure * Must have 4+ years with CI/CD and automation tools ... and SRE use cases * Must have proficiency to debug or troubleshoot and/or deploying SQL and/or ...

SRE Engineer

Washington, DC · On-site

$64.50 - $85.75/hr

Job Title: SRE Engineer Location: Washington, DC Duration: 12+ Months Rate: DOE Reliability ... Amazon ECS. * Automation & Scripting: Mid-level proficiency in Python (or similar scripting ...

SRE Engineer (Technical Lead)

Austin, TX · On-site

$56.50 - $75/hr

Required : • Amazon Web Services (AWS) • AWS IAM • Ansible • Splunk • DevOps / SRE experience Company : Founded and incorporated in 2012 , Info Way Solutions is an IT services and ...

Site Reliability Engineer (SRE)

Plano, TX · On-site

$54.50 - $72.50/hr

Site Reliability Engineer (SRE) Location: Richmond, VA or Plano, TX Work Model: Hybrid - 3 days onsite per week Duration: Long term contract Job Summary: We are seeking an experienced Site ...

Site Reliability Engineer

Denver, CO · On-site

$250/day

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

As a SRE, you will be responsible for maintaining and improving uptime and availability across ... Linux (CentOS/RedHat/Oracle/Amazon Linux) * Programming languages: Python, JavaScript, Java

next page

Showing results 1-20

Amazon Sre information

See salary details

$10

$63

$91

How much do amazon sre jobs pay per hour?

As of Aug 13, 2026, the average hourly pay for amazon sre in the United States is $63.74, according to ZipRecruiter salary data. Most workers in this role earn between $54.81 and $72.84 per hour, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive in the Amazon SRE position, and why are they important?

To thrive as an Amazon SRE (Site Reliability Engineer), you need a strong background in computer science, systems engineering, and automation, typically demonstrated through a relevant degree or equivalent experience. Familiarity with cloud services (especially AWS), containerization tools (like Docker or Kubernetes), monitoring solutions, and infrastructure-as-code platforms, along with certifications (such as AWS Certified DevOps Engineer), is highly valued. Problem-solving mindset, effective communication, and the ability to collaborate across multidisciplinary teams are essential soft skills. These competencies help maintain high system reliability, enable efficient incident response, and support the rapid innovation environment at Amazon.

Are Amazon SRE roles in demand?

Amazon SRE roles are in high demand due to the company's reliance on scalable, reliable cloud infrastructure. These positions often require expertise in cloud platforms, automation, and monitoring tools, and are expected to grow as organizations prioritize system reliability and operational efficiency.

What are the typical daily responsibilities of an Amazon SRE, and how does the role interact with other teams?

As an Amazon SRE, your daily responsibilities include monitoring system performance, proactively identifying and resolving reliability issues, automating repetitive tasks, and participating in on-call rotations to address production incidents. You will also work on infrastructure improvements, conduct root cause analyses after outages, and help implement robust deployment pipelines. Collaboration is frequent, as SREs partner closely with software engineers, product managers, and operations staff to ensure applications run smoothly and reliably. This teamwork fosters a culture of continuous improvement, making Amazon's large-scale services resilient and highly available.

What is an Amazon SRE?

An Amazon Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of Amazon's infrastructure and services. They bridge the gap between development and operations, applying software engineering principles to system administration tasks. SREs focus on automation, monitoring, and incident response to minimize downtime and improve efficiency. Their goal is to create highly available and resilient systems while balancing innovation and system stability.

More about Amazon Sre jobs
What cities are hiring for Amazon Sre jobs? Cities with the most Amazon Sre job openings:
What states have the most Amazon Sre jobs? States with the most job openings for Amazon Sre jobs include:
Infographic showing various Amazon Sre job openings in the United States as of August 2026, with employment types broken down into 64% Full Time, and 36% Contract. Highlights an 93% In-person, and 7% Remote job distribution, with an average salary of $132,583 per year, or $63.7 per hour.

Site Reliability Engineer (SRE)

ATLANTICUS

Atlanta, GA • On-site

$54.75 - $72.75/hr

Full-time

Medical, Retirement, PTO

Posted 5 days ago


Job description

Site Reliability Engineer (SRE)

When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage entrepreneurial thinking to empower our customers toward financial well-being. 

Atlanticus™ technology enables bank, retail, and healthcare partners to offer more inclusive financial services to everyday Americans through the use of proprietary analytics. We apply the experience gained and infrastructure built from servicing over 20 million customers and over $40 billion in consumer loans over more than 25 years of operating history to support lenders that originate a range of consumer loan products. These products include retail and healthcare, private label credit and general-purpose credit cards marketed through our omnichannel platform, including retail point-of-sale, healthcare point-of-care, direct mail solicitation, digital marketing, and partnerships with third parties. Additionally, through our Auto Finance subsidiary, Atlanticus serves the individual needs of automotive dealers and automotive non-prime financial organizations with multiple financing and service programs. 

Office Locations available for this role include 

  • Austin, TX – Situated in The Domain, a vibrant tech hub with park-like surroundings, top restaurants, and convenient parking, perfect for post-work socializing. 
  • Atlanta, GA – Located in the Queen Building (King & Queen Towers, Sandy Springs), with easy access to I-285, GA-400, and a free shuttle to MARTA. 

Work Culture 

We foster a collaborative, innovative environment where everyone contributes to building something meaningful. You’ll be empowered to lead, grow, and make an impact. 

The Role 

We are seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational excellence of our cloud-native applications running on AWS. This is a hands-on role responsible for monitoring and supporting production systems, automating operational tasks, managing deployments, and driving continuous improvements in system stability.

The ideal candidate has strong experience supporting Java-based applications running on Amazon EKS, a solid understanding of AWS infrastructure, and expertise with observability platforms such as Datadog and Splunk. This role requires participation in a 24x7 production support and on-call rotation, working closely with Development, DevOps, IT Ops, Database, Network, and Security teams to maintain highly available production services.

The successful candidate should be passionate about automation, troubleshooting complex production issues, improving application reliability, and leveraging AI-powered tools to enhance operational efficiency.

Key Responsibilities

  • Provide 24x7 production support through an on-call rotation to ensure application availability and rapid incident response.
  • Continuously monitor production applications, infrastructure, and platform health using Datadog, Splunk, CloudWatch, and other monitoring tools.
  • Respond to production incidents, troubleshoot issues, and restore services while minimizing customer impact.
  • Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
  • Deploy and support Java-based applications running on Docker and Amazon EKS using CI/CD pipelines.
  • Execute production deployments, application releases, hotfixes, and rollbacks following change management processes.
  • Monitor and manage scheduled application jobs, batch processes, and integrations to ensure successful execution.
  • Troubleshoot Java application issues using logs, JVM metrics, thread dumps, heap dumps, and application performance metrics.
  • Analyze application, infrastructure, and Kubernetes logs using Splunk and Datadog to identify performance bottlenecks and operational issues.
  • Develop automation scripts using Python, Bash, or similar scripting languages to eliminate repetitive operational tasks.
  • Build self-healing and automated operational processes to improve system reliability and reduce manual intervention.
  • Support Kubernetes (Amazon EKS) environments, including troubleshooting pods, deployments, networking, ingress, and scaling issues.
  • Maintain and improve dashboards, alerts, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational runbooks.
  • Partner with Development teams to improve application reliability, resiliency, scalability, and performance.
  • Continuously improve operational processes, monitoring coverage, automation, and deployment practices.
  • Leverage AI-powered engineering tools and agentic AI capabilities to improve monitoring, incident response, automation, and operational efficiency.

You’re a great fit if you have 

  • 5+ years of experience supporting production applications in an SRE, DevOps, Production Support, or Site Reliability Engineering role.
  • Strong experience supporting Java-based applications in production environments.
  • Hands-on experience with AWS services including EKS, EC2, ALB/NLB, RDS, IAM, Route 53, CloudWatch, S3, and VPC.
  • Experience with Kubernetes (Amazon EKS), Docker, and containerized application deployments.
  • Strong experience using Datadog / Splunk for infrastructure monitoring, APM, troubleshooting, dashboards, alerting, and log analysis.
  • Experience performing production deployments through CI/CD pipelines (Jenkins, GitHub Actions, Argo CD, or similar).
  • Experience supporting MySQL and Oracle databases from an application support perspective.
  • Proficiency in Python, Bash, or other scripting languages for automation.
  • Strong Linux system administration and troubleshooting skills.
  • Excellent troubleshooting skills across distributed applications, networking, and cloud infrastructure.
  • Knowledge of networking fundamentals including DNS, TCP/IP, HTTP/HTTPS, TLS, load balancing, and firewalls.
  • Experience with incident management, problem management, and change management processes.
  • Experience using AI-assisted development tools or agentic AI systems to improve operational efficiency.

Preferred  

  • Experience with Helm and GitOps deployment models.
  • Experience with Terraform or Infrastructure as Code.
  • Familiarity with Prometheus, Grafana, or OpenTelemetry.
  • Experience supporting microservices architectures.
  • Knowledge of JVM tuning and Java performance optimization.
  • Experience with AWS Auto Scaling, Karpenter, or Cluster Autoscaler.

Why You’ll Love Working Here 

This isn’t just a job, it’s a place to lead, grow, and thrive. If you believe in your skills and drive, we’ll provide the resources and support to help you succeed. 

Benefits include 

  • Generous PTO and holiday schedule 
  • 401(k) with company match 
  • Employee stock purchase plan 
  • Ongoing training (lunch & learns, financial and health webinars) 
  • Team volunteer outings 

Atlanticus is an equal opportunity employer. All qualified applicants will receive consideration without regard to race, religion, gender, sexual orientation, age, veteran status, disability, or other protected status.