1

Site Reliability Engineering Jobs in Michigan (NOW HIRING)

Senior Systems Engineer

Southfield, MI · Hybrid

$95K - $131K/yr

Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems management. * Experience with TLS, digital certificates, and related troubleshooting. * Familiarity ...

Senior Systems Engineer

Southfield, MI · Hybrid

$95K - $131K/yr

Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems management. * Experience with TLS, digital certificates, and related troubleshooting. * Familiarity ...

The role leads a multidisciplinary technical team spanning Platform Engineering, Data Engineering, Data Product Enablement, and Site Reliability Engineering. The Platform Lead sets technical ...

Reliability Engineer

Battle Creek, MI · On-site

$92K - $116K/yr

Bachelor's Degree in Engineering preferred, but we're casting a wide net for the right problem ... This role is designed for on-site partnership with maintenance and operations teams, with daily ...

Bachelor's Degree in Engineering preferred, but we're casting a wide net for the right problem ... This role is designed for on-site partnership with maintenance and operations teams, with daily ...

Sr Software Engineer TECHM-JOB-22848

Lansing, MI · On-site

$111K - $147K/yr

... Site Reliability Engineering (SRE) and Infrastructure as Code. • Extensive knowledge of Storage Area Networks (SAN) based servers. • Project management experience including driving projects ...

Showing results 41-60

Site Reliability Engineering information

See Michigan salary details

$9

$55

$80

How much do site reliability engineering jobs pay per hour?

As of Sep 6, 2026, the average hourly pay for site reliability engineering in Michigan is $55.56, according to ZipRecruiter salary data. Most workers in this role earn between $47.79 and $63.46 per hour, depending on experience, location, and employer.

What is site reliability engineering?

Site Reliability Engineering (SRE) is a discipline that combines aspects of software engineering and IT operations to build and run scalable, reliable, and efficient systems. SRE teams are responsible for ensuring the availability, performance, and reliability of critical services by automating manual processes, monitoring systems, and responding to incidents. They often work closely with development teams to improve system architecture, deploy new features safely, and maintain service-level objectives (SLOs). Overall, SRE aims to create a bridge between development and operations to deliver robust and reliable software services.

How does a site reliability engineer typically collaborate with development and operations teams?

Site Reliability Engineers (SREs) work closely with both development and operations teams to ensure systems are reliable, scalable, and efficient. They often participate in code reviews, help define service level objectives (SLOs), and develop automation tools to streamline deployment and incident response. SREs act as a bridge between development and IT, translating operational needs into engineering solutions and vice versa. Regular communication and joint problem-solving are essential parts of the role, fostering a culture of shared responsibility for system uptime and performance.

What are the key skills and qualifications needed to thrive as a site reliability engineer, and why are they important?

To thrive as a Site Reliability Engineer, you need a solid background in software engineering, systems administration, and troubleshooting, often supported by a degree in computer science or related field. Familiarity with automation tools, cloud platforms (such as AWS, GCP, or Azure), containerization (Docker, Kubernetes), and monitoring systems is typically required. Strong problem-solving skills, effective communication, and a proactive mindset set outstanding SREs apart. These skills ensure high system reliability, efficient incident response, and seamless collaboration across development and operations teams.

What are the most commonly searched types of Site Reliability Engineering jobs in Michigan?

The most popular types of Site Reliability Engineering jobs in Michigan are:

Infographic showing various Site Reliability Engineering job openings in Michigan as of August 2026, with employment types broken down into 1% As Needed, 77% Full Time, 20% Part Time, and 2% Contract. Highlights an 93% Physical, 3% Hybrid, and 4% Remote job distribution, with an average salary of $115,559 per year, or $55.6 per hour.

Senior Systems Engineer

CARFAX

Southfield, MI • Hybrid

$95K - $131K/yr

Full-time

Re-posted 16 hours ago


Key responsibilities

  • Build, install, and administer on-premises systems, servers, hardware, and storage in accordance with security standards and operational requirements.

  • Architect, integrate, and administer enterprise cloud and hybrid infrastructure solutions, ensuring alignment with business requirements for reliability and disaster recovery.

  • Build and maintain monitoring systems and dashboards; develop tools and scripts to automate and streamline server and container administration.


Job description

Join our winning team as a Senior Systems Engineer
 
We are seeking a skilled Systems Engineer to join our Enterprise Systems Engineering team. The team is highly skilled and collaborative, diverse, and geographically distributed. Our mission is to build, administer, secure, and continuously improve thousands of servers and systems across the entire enterprise while delivering new solutions that align with strategic business objectives.
 
The ideal candidate will demonstrate a proven track record in large-scale server administration, possessing broad expertise in configuring and managing on-premises and multi-cloud environments. Experience migrating on-prem systems to the cloud, with a strong hands-on background with containerization, Infrastructure as Code, automation, and networking are essential.
 
This is a hybrid role with in-office expectations determined by location and subject to change based on evolving business needs. 
 
What you'll be doing:
  • Build, install, and administer on-premises systems, servers, hardware, and storage in accordance with security standards and operational requirements.
  • Architect, integrate, and administer enterprise cloud and hybrid infrastructure solutions, ensuring alignment with business requirements for reliability and disaster recovery. 
  • Establish and promote best practices, standards, training, and general support for both on-premises and cloud environments.
  • Continuously analyze and develop solutions to improve platform performance, reliability, scalability, and cost efficiency across applications and processes. 
  • Lead design decisions, including researching and prototyping emerging technologies.
  • Champion exceptional containerization and cloud platform design and quality.
  • Build and maintain monitoring systems and dashboards; develop tools and scripts to automate and streamline server and container administration.
  • Provide operational support through on-call rotations, ensuring high availability and timely incident resolution, including root cause analysis.
          What we're looking for:
          • 7+ years of experience building, hardening, and administering Linux in a production environment. 
          • 5+ years of hands-on experience with Kubernetes clusters, including setup, configuration, monitoring, and troubleshooting (AWS EKS experience is a plus). 
          • 5+ years of experience designing, supporting, and migrating on-premises resources to AWS cloud infrastructure, with hands-on experience across core services including EC2, VPC, API Gateway, and related CLI and SDK tooling. 
          • 3+ years of hands-on experience with configuration management and Infrastructure as Code tools such as Chef, Ansible, Puppet, Terraform, and AWS CDK. 
          • Proficiency in Python, YAML, Bash, PowerShell, or other scripting and object-oriented languages. 
          • Broad expertise in DevSecOps, encompassing security and performance best practices for on-premises, cloud, and hybrid architectures, with a demonstrated commitment to continuous learning, documentation, and mentoring. 
          • Operational ownership of system uptime and quality, including advanced troubleshooting and root cause analysis across diverse systems and technologies. 
          • Experience with observability tools such as Prometheus, Grafana, and New Relic. 
          • Experience with Identity and Access Management (IAM) and Privileged Access Management (PAM) solutions. 
          Preferred Qualifications
          • Multi-cloud experience spanning AWS, GCP, and Azure.
          • Familiarity with F5, Gloo, Lambda, VPC, load balancing, and diagnosing server- and network-related issues.
          • Site Reliability Engineering (SRE) experience.
          • Experience with Active Directory and file systems management.
          • Experience with TLS, digital certificates, and related troubleshooting.
          • Familiarity with CloudWatch, Splunk, Logstash, or other log aggregation tools.