1

Customer Reliability Engineer Jobs in Chicago, IL

Staff SRE - Observability

Chicago, IL

$58.75 - $78/hr

We are experts in product practices but life long learners in the domain of our customers. We ... Experience implementing SRE practices including error budgets and toil metrics * Proficiency in ...

Site Reliability Engineer, Observability

Chicago, IL · On-site

$58.75 - $78/hr

... customers. The incident management program you will help build is early-stage-you will be ... Core SRE Experience * 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering ...

SRE Cloud Engineer

Chicago, IL · On-site

$58.75 - $78/hr

Financial Services Domain , predominantly Wealth management • Technical : SRE Tools ... customer experience • Communicate issue/resolution status (written and verbal) to project team ...

Showing results 41-60

Customer Reliability Engineer information

See Chicago, IL salary details

$62.8K

$121.5K

$145.3K

How much do customer reliability engineer jobs pay per year?

As of Aug 8, 2026, the average yearly pay for customer reliability engineer in Chicago, IL is $121,529.00, according to ZipRecruiter salary data. Most workers in this role earn between $105,600.00 and $132,900.00 per year, depending on experience, location, and employer.

How does a customer reliability engineer typically interact with clients and internal engineering teams?

Customer Reliability Engineers serve as a vital bridge between clients and internal technical teams. They regularly communicate with customers to understand their needs, troubleshoot issues, and provide technical guidance. Internally, they collaborate closely with product, support, and development teams to relay customer feedback, help prioritize reliability improvements, and ensure seamless incident resolution. This cross-functional role requires strong communication skills and the ability to translate technical information for different audiences, making every day varied and impactful.

What does a customer reliability engineer do?

A customer reliability engineer (CRE) works with clients to ensure the reliability, performance, and availability of products or services. They analyze system issues, develop solutions, and often collaborate with engineering teams to improve infrastructure and customer experience, typically using monitoring tools and technical expertise. CREs may also provide technical support and guidance to help clients optimize their use of the company's offerings.

What is the difference between Customer Reliability Engineer vs Site Reliability Engineer?

AspectCustomer Reliability EngineerSite Reliability Engineer
CredentialsTypically requires engineering degrees, certifications in cloud platforms (AWS, Azure), and knowledge of customer supportRequires engineering degrees, certifications in cloud and systems management, with a focus on infrastructure
Work EnvironmentCustomer-facing, involves direct interaction with clients to resolve issues and improve reliabilityPrimarily internal, focused on maintaining and improving system reliability and scalability
Employer & Industry UsageUsed by cloud service providers and tech companies with a customer support componentCommon in large tech companies managing large-scale infrastructure and services

The main difference is that Customer Reliability Engineers focus on ensuring customer satisfaction and resolving client-specific issues, while Site Reliability Engineers concentrate on internal system stability and scalability. Both roles require technical expertise and cloud knowledge but serve different operational needs.

What is a customer reliability engineer?

A Customer Reliability Engineer (CRE) is a technical professional who works closely with customers to ensure the reliability, performance, and uptime of software products and services. CREs act as a bridge between customers and engineering teams, helping to identify, troubleshoot, and resolve reliability issues. They often collaborate with multiple departments to implement best practices, monitor systems, and proactively address potential problems, ultimately aiming to improve the overall customer experience.

What skills and qualifications are needed to thrive as a customer reliability engineer?

To thrive as a Customer Reliability Engineer, you need a solid background in systems engineering, incident management, and troubleshooting, often supported by a degree in computer science or related field. Familiarity with cloud platforms (such as AWS or GCP), monitoring tools (like Datadog or Prometheus), and automation scripts is typically required. Exceptional communication, problem-solving abilities, and a customer-centric mindset are vital soft skills for this role. These skills ensure efficient incident resolution, strong client relationships, and reliable system performance under pressure.
What are popular job titles related to Customer Reliability Engineer jobs in Chicago, IL? For Customer Reliability Engineer jobs in Chicago, IL, the most frequently searched job titles are:
What job categories do people searching Customer Reliability Engineer jobs in Chicago, IL look for? The top searched job categories for Customer Reliability Engineer jobs in Chicago, IL are:
What cities near Chicago, IL are hiring for Customer Reliability Engineer jobs? Cities near Chicago, IL with the most Customer Reliability Engineer job openings:

Staff SRE - Observability

Focused

Chicago, IL

$58.75 - $78/hr

Full-time

Re-posted 17 days ago


Job description

Who we are:

At Focused, we move quickly to deliver quality software that achieves client outcomes and meets their customer's needs. We strategically partner with our clients to leverage our expertise in design and software, while our clients bring their own domain expertise. We work with a variety of clients from different industries, collaborating as we get new products to market, modernizing legacy systems, or helping teams learn the skills they need to be successful.

Our values:

  • Listen first • We are experts in product practices but life long learners in the domain of our customers. We research, collaborate, and understand.
  • Learn why • We ask questions and talk to users to understand problem spaces, objectives, and goals, which allows us to deeply invest and drive towards the outcomes of our clients.
  • Love your craft • We love diving into a variety of domains and solving problems. We take pride in delivering value, in communicating progress, and guiding our clients to success.

We are seeking an experienced Staff Observability Consultant with deep expertise in OpenTelemetry and strong Platform Engineering capabilities to help organizations implement, optimize, and scale their observability infrastructure. This role requires a seasoned consultant who can design comprehensive telemetry strategies, implement distributed tracing solutions, establish robust monitoring practices, and interface closely with clients on the observability journey.

Key Responsibilities:

OpenTelemetry & Observability

  • Design and implement end-to-end OpenTelemetry solutions across diverse technology stacks
  • Configure and deploy OpenTelemetry Collectors for efficient data collection, processing, sampling, and routing
  • Establish telemetry pipelines for metrics, traces, and logs across microservices architectures
  • Optimize collector configurations for performance, reliability, and cost-effectiveness

Platform Engineering & Infrastructure

  • Augment existing infrastructure with with integrated observability solutions
  • Implement Infrastructure as Code (IaC) solutions using Terraform, Pulumi, CloudFormation, etc.
  • Architect and manage Kubernetes clusters with comprehensive monitoring and logging
  • Build CI/CD pipelines with embedded observability and automated testing

Site Reliability Engineering (SRE)

  • Establish and maintain Service Level Indicators (SLIs), Objectives (SLOs), and Agreements (SLAs)
  • Implement error budgets, toil reduction strategies, and capacity planning
  • Support incident response procedures and post-mortem processes

Cloud & DevOps Engineering

  • Deploy and manage observability infrastructure across AWS, GCP, and Azure
  • Establish security, compliance, and governance frameworks for telemetry data
  • Experience automating Agent Evaluations in CI/CD pipelines and observability backends.

Required Qualifications:

Core Observability & OpenTelemetry

  • 3-7 years of experience in observability, monitoring, and distributed systems
  • Deep hands-on experience with OpenTelemetry ecosystem, including SDKs, APIs, and specifications
  • Proficiency with OpenTelemetry Collector configuration, processors, exporters, and receivers
  • Strong understanding of telemetry data models, semantic conventions, and instrumentation best practices

Platform Engineering & DevOps

  • 5+ years of Platform Engineering or DevOps experience with focus on site reliability, observability, and incident response
  • Proficiency with Infrastructure as Code tools (Terraform, Pulumi, CloudFormation, CDK)
  • Strong experience with CI/CD platforms (GitHub Actions, GitLab CI, Jenkins, ArgoCD)

Cloud & Infrastructure

  • Hands-on experience with major cloud providers (AWS, GCP, Azure) and their observability services
  • Experience with container technologies (Docker, Podman) and container registries
  • Knowledge of networking, security, load balancing, and distributed systems concepts

Site Reliability Engineering

  • Experience implementing SRE practices including error budgets and toil metrics
  • Proficiency in incident management, on-call procedures, and post-mortem culture
  • Experience with capacity planning, performance optimization, and scalability design

Programming & Automation

  • Proficiency in multiple programming languages preferred (Go, Python, Java, Node.js, Rust)
  • Strong scripting and automation skills (Bash, Python, PowerShell)
  • Understanding of software engineering best practices and testing methodologies

Preferred Qualifications (Exceptional Candidates)

AI & Agentic Frameworks

  • Understanding of Large Language Models (LLMs) and their application in DevOps
  • Knowledge of vector databases, embeddings, and retrieval-augmented generation (RAG)
  • Experience with AI/ML model deployment and monitoring in production environments

Leadership & Communication

  • Strong technical writing and documentation skills
  • Ability to present complex technical concepts to diverse stakeholders
  • A passion for knowledge sharing

Key Competencies

  • Systems thinking and ability to design holistic observability solutions
  • Strong analytical and troubleshooting skills for complex distributed systems
  • Curiosity about emerging technologies, particularly AI applications in operations
  • Adaptability to rapidly evolving cloud-native and observability technologies
  • Collaborative mindset with focus on enabling developer productivity and system reliability

What Sets Exceptional Candidates Apart:

  • Experience with Honeycomb
  • Contributions to open-source observability or AI framework projects
  • Track record of implementing platform engineering solutions that significantly improved developer experience
  • Experience scaling observability infrastructure to handle high event volume

What to know before you apply:

  • This role will require being in the Chicago office three days per week and up to 20% travel within the United States.
  • Focused is unable to sponsor or take over sponsorship of the employment Visa process at this time.
  • The Chicago base salary range for this role is $160,000 - $200,000.