1

Cloud Reliability Analyst Jobs (NOW HIRING)

Platform & Cloud Reliability (AWS, GCP, Snowflake, EMR, Hadoop, ETL/ELT) * Leverage Enterprise ... Champion the adoption of machine learning-based observability and reliability analytics. EndtoEnd ...

Lead cloud reliability engineering (SRE), DevOps, and platform automation practices. * Evaluate ... Advanced analytical, architectural, and problem-solving abilities. Experience: * 12+ years of ...

Lead cloud reliability engineering (SRE), DevOps, and platform automation practices. * Evaluate ... Advanced analytical, architectural, and problem‑solving abilities. Experience * 12+ years of ...

Showing results 21-40

Cloud Reliability Analyst information

See salary details

$38.5K

$91.9K

$140K

How much do cloud reliability analyst jobs pay per year?

As of Sep 9, 2026, the average yearly pay for cloud reliability analyst in the United States is $91,938.00, according to ZipRecruiter salary data. Most workers in this role earn between $77,000.00 and $104,000.00 per year, depending on experience, location, and employer.

What is a cloud reliability analyst?

A Cloud Reliability Analyst is a professional responsible for ensuring the reliability, availability, and performance of cloud-based systems and services. They monitor cloud infrastructure, troubleshoot incidents, and implement best practices to prevent downtime. Their role also includes collaborating with development and operations teams to optimize system reliability and automate processes. Ultimately, they help organizations maintain seamless and resilient cloud environments.

What are some common challenges cloud reliability analysts face when supporting large-scale cloud environments?

Cloud Reliability Analysts often encounter challenges such as managing complex, distributed systems with varying workloads, ensuring high availability, and quickly identifying the root causes of outages or performance issues. They must stay ahead of evolving security threats and adapt to rapidly changing cloud technologies. Collaborating closely with development, operations, and security teams is essential to proactively address incidents and automate monitoring and recovery processes.

What are the key skills and qualifications needed to thrive as a cloud reliability analyst, and why are they important?

To thrive as a Cloud Reliability Analyst, you need expertise in cloud computing concepts, infrastructure monitoring, incident response, and a related IT or computer science degree. Familiarity with cloud platforms (such as AWS, Azure, or Google Cloud), automation tools, and monitoring systems (like Datadog, Prometheus, or Splunk) is typically required. Strong analytical thinking, problem-solving, and clear communication skills help you proactively address reliability issues and collaborate across teams. These skills are crucial to ensure optimal cloud performance, minimize downtime, and maintain robust, scalable systems for business continuity.

What are popular job titles related to Cloud Reliability Analyst jobs?

For Cloud Reliability Analyst jobs, the most frequently searched job titles are:

Cloud Infrastructure Site Reliability Engineer

Berkeley Heights, NJ • On-site

Judge Group, Inc.
Recruiting and Staffing Services • 5 - 10K employees

$70 - $80/hr

Other

Posted 28 days ago


Key responsibilities

  • Design, build, and maintain highly available, scalable, and secure cloud infrastructure on platforms such as AWS, Google Cloud Platform, or Azure.

  • Develop and implement automation for provisioning, monitoring, scaling, and incident response using Infrastructure-as-Code tools.

  • Monitor system reliability, capacity, and performance; proactively detect and address issues before they impact users.


Job description

Location: Berkeley Heights, NJ Salary: $70.00 USD Hourly - $80.00 USD Hourly Description:
Job Title: Cloud Infrastructure Site Reliability Engineer
Location: Berkeley Heights, NJ / Alpharetta, GA (Onsite 5 Days)
Duration: Contract To Hire
Job Description:
Position Summary:
As a Cloud Infrastructure Site Reliability Engineer (SRE) with expertise in multiple public cloud service provider platforms, you will be responsible for operating infrastructure solutions, following the principles and practices pioneered by Google's SRE model. Your work will ensure our cloud services meet uptime, reliability, and performance targets, and you will drive automation and continuous improvement across our production environments. This role will involve collaborating with cross-functional teams to enhance our cloud reliability posture and streamline processes through automation.
Key Responsibilities:
  • Design, build, and maintain highly available, scalable, and secure cloud infrastructure on platforms such as AWS, Google Cloud Platform, or Azure.
  • Develop and implement automation for provisioning, monitoring, scaling, and incident response using Infrastructure-as-Code tools (e.g., Terraform, CloudFormation, Ansible).
  • Monitor system reliability, capacity, and performance; proactively detect and address issues before they impact users.
  • Respond to production incidents, participate in on-call rotations, and lead post-incident reviews to drive root cause analysis and reliability improvements.
  • Collaborate with software engineering and security teams to ensure new services and features are production-ready and meet reliability standards.
  • Build and maintain tools for deployment, monitoring, and operations; automate manual processes to reduce toil.
  • Document operational processes and system architectures to ensure knowledge sharing and repeatability.
  • Continuously evaluate and implement new technologies to improve system reliability, security, and efficiency.

Qualifications:
  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 3+ years of experience in software development with proficiency in at least one programming language (e.g., Python, Go, Java, C++).
  • Experience administrating cloud platforms (AWS, Google Cloud Platform, Azure), including networking, security, containerization, storage, data management, and serverless technologies.
  • Solid understanding of Linux systems, networking fundamentals, virtualized, and distributed systems, file systems, system processes and configurations.
  • Deep understanding of observability (monitoring, alerting, and logging) tools in cloud environments. Ability to set up and maintain monitoring dashboards, alerts, and logs.
  • Familiarity with Continuous Integration/Continuous Deployment (CI/CD) tools for automated testing, deployments, provisioning, and observability.
  • Ability to manage and respond to incidents, perform root cause analysis, and implement post-mortem reviews.
  • Understanding of setting, monitoring, and maintaining Service-Level Objectives (SLOs) and Service-Level Agreements (SLAs) for system reliability.

Needs experience with Terraform and Dynatrace
  • Additional Qualifications a Plus: Experience working with enterprise-scale financial services or other regulated industries
  • 5+ years of experience in SRE, DevOps, infrastructure, or cloud engineering roles, preferably supporting large-scale, distributed systems.
  • Excellent problem-solving, troubleshooting, and communication skills.
  • Experience leading technical projects or mentoring junior engineers.
  • Certifications: Certified Engineer, DevOps, SRE, CSREF

By providing your phone number, you consent to: (1) receive automated text messages and calls from the Judge Group, Inc. and its affiliates (collectively "Judge") to such phone number regarding job opportunities, your job application, and for other related purposes. Message & data rates apply and message frequency may vary. Consistent with Judge's Privacy Policy, information obtained from your consent will not be shared with third parties for marketing/promotional purposes. Reply STOP to opt out of receiving telephone calls and text messages from Judge and HELP for help.
Contact:
This job and many more are available through The Judge Group. Please apply with us today!