1

Senior Site Reliability Engineer Jobs in Colorado

Site Reliability Engineer (SRE)

Boulder, CO ยท On-site

$130K - $155K/yr

Amplifire is seeking a Site Reliability Engineer to improve the reliability, scalability, performance, and operational efficiency of our cloud-based platform. Working alongside DevOps engineers ...

Site Reliability Engineer

Aurora, CO ยท On-site

$62 - $141/hr

As a Site Reliability Engineer (SRE) on our team, you'll help the Intelligence Community develop more robust systems by building a resilient infrastructure. You'll build in redundancy, implement ...

CO ยท On-site

You'll collaborate or embed with engineering teams, helping them to improve the reliability and ... Location Find out more about our locations by visiting our site. Compensation & Benefits The base ...

Site Reliability Engineer

Westminster, CO

$57.50 - $76.50/hr

Architect the Future as our Site Reliability Engineer! Are you ready to take your skills to the next level as a self-motivated and enthusiastic Site Reliability Engineer with hands-on experience ...

Site Reliability Engineer

Westminster, CO ยท On-site

$57.50 - $76.50/hr

Architect the Future as our Site Reliability Engineer! Are you ready to take your skills to the next level as a self-motivated and enthusiastic Site Reliability Engineer with hands-on experience ...

Site Reliability Engineer

Westminster, CO ยท On-site

$105K - $145K/yr

Essential Skills & Experience * 3+ years of professional experience in SRE, DevOps, or Cloud Engineering within enterprise software environments. * Azure Mastery: Extensive hands-on experience with ...

Site Reliability Engineer

Westminster, CO ยท On-site

$105K - $145K/yr

Essential Skills & Experience * 3+ years of professional experience in SRE, DevOps, or Cloud Engineering within enterprise software environments. * Azure Mastery: Extensive hands-on experience with ...

Site Reliability Engineer

Westminster, CO ยท On-site

$57.50 - $76.50/hr

Architect the Future as our Site Reliability Engineer! Are you ready to take your skills to the next level as a self-motivated and enthusiastic Site Reliability Engineer with hands-on experience ...

CO ยท On-site

$138K - $173K/yr

You'll collaborate or embed with engineering teams, helping them to improve the reliability and ... Location Find out more about our locations by visiting our site. Compensation & Benefits The base ...

Site Reliability Engineer

Aurora, CO ยท On-site

$62K - $141K/yr

Site Reliability Engineer The Opportunity: Engineering to make a system more resilient and efficient frees up time and money to build more capabilities. Whether you come from a background in network ...

Staff Site Reliability Engineer

Greenwood Village, CO ยท On-site

$57.75 - $76.75/hr

As a Staff Site Reliability Engineer at TransUnion, you will serve as a senior technical leader and force multiplier on the SRE team. Operating with full autonomy, you will drive reliability strategy ...

Staff Site Reliability Engineer

Greenwood Village, CO ยท On-site

$57.75 - $76.75/hr

As a Staff Site Reliability Engineer at TransUnion, you will serve as a senior technical leader and force multiplier on the SRE team. Operating with full autonomy, you will drive reliability strategy ...

Site Reliability Engineer

Westminster, CO ยท On-site +1

$105K - $145K/yr

Trimble is seeking a Site Reliability Engineer to join our globally diverse team helping develop and support our FedRAMP PaaS portfolio. This role is a key member of cloud-hosted solutions residing ...

Site Reliability Engineer

Westminster, CO ยท On-site

$91K - $125K/yr

Elevate Scalability as Our Next Site Reliability Engineer! Ready to make a high-impact mark on critical infrastructure using AWS GovCloud and cutting-edge cloud technologies? Trimble is looking for ...

Showing results 21-40

Senior Site Reliability Engineer information

See Colorado salary details

$22

$67

$97

How much do senior site reliability engineer jobs pay per hour?

As of Sep 3, 2026, the average hourly pay for senior site reliability engineer in Colorado is $67.73, according to ZipRecruiter salary data. Most workers in this role earn between $55.87 and $81.15 per hour, depending on experience, location, and employer.

What is a senior site reliability engineer?

Senior Site Reliability Engineers (SREs) are experienced IT professionals who ensure the reliability, scalability, and performance of complex software systems. They bridge the gap between software development and operations by automating processes, monitoring system health, responding to incidents, and implementing best practices for system stability. Senior SREs often lead teams, design resilient infrastructure, and mentor junior engineers to improve overall reliability and efficiency in technology environments.

What are the key skills and qualifications needed to thrive as a senior site reliability engineer?

To thrive as a Senior Site Reliability Engineer, you need deep experience in systems administration, software engineering, automation, and cloud infrastructure, typically supported by a related degree and several years in SRE or DevOps roles. Proficiency with tools like Kubernetes, Docker, Terraform, monitoring platforms (e.g., Prometheus, Datadog), and certifications in AWS, GCP, or Azure are common requirements. Strong problem-solving skills, effective communication, and a proactive mindset help you collaborate across teams and respond quickly to incidents. These skills ensure high system reliability, scalability, and rapid recovery from failures, which are critical to business continuity.

What are some typical challenges a senior site reliability engineer faces when balancing system reliability with rapid feature releases?

Senior Site Reliability Engineers often navigate the challenge of maintaining high system uptime and stability while supporting fast-paced software development cycles. This involves implementing robust automation and monitoring, collaborating closely with development teams to ensure new features don't compromise reliability, and proactively addressing potential bottlenecks. Effective communication and prioritization are key, as SREs must advocate for best practices in reliability without slowing down innovation. Continuous learning and adaptation are vital to keep up with evolving technology stacks and complex distributed systems.

What does a senior site reliability engineer do?

A senior site reliability engineer (SRE) is responsible for maintaining and improving the reliability, availability, and performance of large-scale systems and services. They develop automation tools, monitor system health, troubleshoot issues, and implement best practices to ensure continuous operation, often using skills in coding, systems administration, and cloud platforms.

What are the most commonly searched types of Site Reliability Engineer jobs in Colorado?

The most popular types of Site Reliability Engineer jobs in Colorado are:

What are popular job titles related to Senior Site Reliability Engineer jobs in Colorado?

For Senior Site Reliability Engineer jobs in Colorado, the most frequently searched job titles are:

What cities in Colorado are hiring for Senior Site Reliability Engineer jobs?

Cities in Colorado with the most Senior Site Reliability Engineer job openings:

Infographic showing various Senior Site Reliability Engineer job openings in Colorado as of August 2026, with employment types broken down into 1% As Needed, 83% Full Time, 13% Part Time, 1% Temporary, and 2% Contract. Highlights an 93% Physical, 3% Hybrid, and 4% Remote job distribution, with an average salary of $140,883 per year, or $67.7 per hour.

Site Reliability Engineer (SRE)

Amplifire

Boulder, CO โ€ข On-site

$130K - $155K/yr

Full-time

Posted 12 days ago


Job description

Description:

Amplifire is seeking a Site Reliability Engineer to improve the reliability, scalability, performance, and operational efficiency of our cloud-based platform. Working alongside DevOps engineers within the Platform Operations team, this role combines software engineering and systems operations with a focus on observability, automation, incident reduction, and operational excellence.

You will partner closely with software engineering, QA, security, and DevOps to establish reliability practices, improve production visibility, strengthen incident response, automate operational workflows, and help engineering teams deliver changes safely and confidently.
 

This role supports systems operating under regulatory and compliance requirements, including FedRAMP and SOC 2. The ideal candidate understands that reliability, traceability, security, and change management enable sustainable development velocity rather than compete with it.


Amplifire expects all technical team members to leverage AI-assisted tools and workflows as force multipliers for productivity, learning, automation, and problem-solving while maintaining strong engineering skills, sound judgment, security standards, and operational accountability.

Reliability, Observability & Performance

· Establish and maintain service-level indicators (SLIs), service-level objectives (SLOs), error budgets, and other measures of system health.

· Build and continuously improve monitoring, logging, tracing, dashboards, and alerting that provide actionable visibility into application and infrastructure health.

· Analyze system behavior, performance, capacity, and reliability trends to proactively identify risks and improvement opportunities.

· Partner with engineering teams to define reliability requirements and improve the availability, scalability, and performance of production systems.


Incident Response & Operational Readiness

· Participate in the on-call rotation and respond to production incidents with urgency and sound technical judgment.

· Improve incident detection, triage, escalation, communication, mitigation, and recovery processes.

· Lead or contribute to blameless post-incident reviews, root cause analysis, and corrective actions.

· Develop and maintain runbooks, recovery procedures, and operational practices that improve service resilience and reduce mean time to detect and restore service.


Automation, Infrastructure & Delivery

· Identify and eliminate operational toil through automation, self-service tooling, and continuous improvement.

· Build and maintain cloud infrastructure using Infrastructure as Code practices and tools such as Terraform or AWS CDK.

· Improve CI/CD pipelines, deployment safeguards, rollback capabilities, and progressive delivery practices.

· Develop internal tools and automation that improve reliability, resilience, and engineering productivity.

· Support scalable, secure, and cost-effective production environments.


Collaboration & Continuous Improvement

· Partner with software engineers to embed reliability, operability, and observability throughout the development lifecycle while reducing operational friction and helping teams safely own their services in production.

· Help engineering teams diagnose complex production issues across application and infrastructure layers.

· Contribute to security, compliance, capacity-planning, cloud cost optimization, and operational standards.

· Use AI-assisted tools responsibly to improve troubleshooting, automation, documentation, and operational efficiency while sharing knowledge and continuously improving team practices.


Requirements:

· 4+ years of experience in Site Reliability Engineering, DevOps, cloud infrastructure, systems engineering, software engineering, or a related production-focused role.

· Demonstrated experience supporting reliable, customer-facing applications in a production cloud environment.

· Hands-on experience designing, deploying, and operating production workloads on AWS.

· Experience building or maintaining Infrastructure as Code, observability solutions, incident response processes, and CI/CD automation.

· Experience identifying and reducing operational toil through automation and continuous improvement.

· Experience participating in on-call rotations, troubleshooting production issues, performing root cause analysis, and implementing preventive improvements.


Technical Skills

· Proficiency with at least one scripting or programming language, such as Python, Bash, JavaScript/TypeScript, Go, or Java.

· Working knowledge of Linux systems, networking, DNS, HTTP, load balancing, and common cloud architecture patterns.

· Experience with containers and container-based deployment practices; Docker experience is required, while Kubernetes or similar orchestration experience is beneficial.

· Ability to analyze logs, metrics, traces, and system behavior to troubleshoot issues across application and infrastructure layers.

· Understanding of reliability concepts such as SLIs, SLOs, error budgets, availability, latency, capacity planning, and graceful degradation.

· Familiarity with secure configuration, secrets management, access controls, vulnerability remediation, and other operational security fundamentals. 


AI & Modern Tooling

· Experience using AI-assisted development or operational tools to improve productivity, automation, troubleshooting, documentation, or engineering workflows.

· Ability to critically evaluate AI-generated outputs and apply appropriate validation before production use.

· Interest in adopting emerging tools and practices that improve engineering effectiveness while maintaining operational excellence.


Collaboration & Problem-Solving Skills

· Strong debugging and systems-thinking skills, including the ability to work through ambiguous, cross-service production issues.

· Ability to communicate clearly during incidents and translate technical findings for engineering and business stakeholders.

· Experience collaborating with software engineering, QA, security, support, and product teams.

· Ability to independently own reliability improvements while seeking input and alignment when appropriate.

· A proactive approach to problem-solving and a willingness to challenge existing practices constructively.


What Success Looks Like

Within your first year:

· Service-level indicators and objectives are established for critical systems.

· Monitoring, alerting, and observability provide actionable visibility across the platform.

· Incident response processes are more structured, repeatable, and measurable.

· Mean time to detect and restore service trends improve through better tooling and operational practices.

· Engineering teams have greater self-service access to operational insights and reliability tooling.

· Manual operational work is reduced through automation and platform improvements.

· Reliability, performance, and scalability risks are identified proactively rather than reactively.

· Demonstrates a proactive approach to problem solving and does not accept existing processes simply because they have historically been done that way.


 Amplifire is the leading AI Learning Platform built on brain science that delivers the proven results of 1:1 expert instruction at enterprise scale. We deliver one-on-one AI instruction at scale, detect where people are confidently wrong, and fix it before it becomes a mistake—reducing training time by 50–80% while driving measurable performance outcomes. Trusted by leading organizations in healthcare, accounting, life sciences and other high-stakes industries, Amplifire enables teams to achieve verified competency faster, reduce risk, and perform at the highest level when it matters most.