1

Site Reliability Engineer Jobs in Detroit, MI (NOW HIRING)

SRE Engineer

Dearborn, MI ยท On-site

$52.75 - $70/hr

As a Site Reliability Engineer at Ford Motor Company, you will play a pivotal role in elevating the performance and dependability of our Marketing and Sales Tech platform and applications. In this ...

Site Reliability Engineer - Networking

Ann Arbor, MI ยท On-site

$55.75 - $74/hr

As a Site Reliability Engineer, you will be focused on supporting a specific, highly available, and very secure production environment. You will analyze the reliability of our environments, use your ...

Site Reliability Engineer II

Detroit, MI

$56.50 - $75/hr

Site Reliability Engineer II The SRE II sits at the intersection of software engineering and platform operations. You will own the reliability, scalability, and operational hygiene of Kastle's core ...

Site Reliability Engineer III

Ann Arbor, MI ยท On-site

$55.75 - $74/hr

Level III Site Reliability Engineers are recognized technical experts who lead complex projects and initiatives, drive innovation, and serve as key resources for both their team and the broader ...

Site Reliability Engineer III

Ann Arbor, MI ยท On-site

$55.75 - $74/hr

Level III Site Reliability Engineers are recognized technical experts who lead complex projects and initiatives, drive innovation, and serve as key resources for both their team and the broader ...

You'll operate at the intersection of classic SRE discipline which include IaC, CI/CD, observability, incident response and the emerging demands of running agentic AI systems in production: LLM ...

Staff Site Reliability Engineer

Ann Arbor, MI ยท On-site +1

$55.75 - $74/hr

You'll operate at the intersection of classic SRE discipline which include IaC, CI/CD, observability, incident response and the emerging demands of running agentic AI systems in production: LLM ...

Cloud Engineer

Dearborn, MI ยท On-site

$51.25 - $68.50/hr

Apply Site Reliability Engineering (SRE) principles to improve availability, reliability, scalability, and operational performance. * Integrate enterprise security, backup, disaster recovery, and ...

Cloud Engineer

Dearborn, MI ยท On-site

$51.25 - $68.50/hr

Apply Site Reliability Engineering (SRE) principles to improve availability, reliability, scalability, and operational performance.Integrate enterprise security, backup, disaster recovery, and data ...

Cloud Engineer

Dearborn, MI ยท On-site

$174.79/hr

Apply Site Reliability Engineering (SRE) principles to improve availability, reliability, scalability, and operational performance. * Integrate enterprise security, backup, disaster recovery, and ...

next page

Showing results 1-20

Site Reliability Engineer information

See Detroit, MI salary details

$9

$58

$84

How much do site reliability engineer jobs pay per hour?

As of Aug 27, 2026, the average hourly pay for site reliability engineer in Detroit, MI is $58.32, according to ZipRecruiter salary data. Most workers in this role earn between $50.14 and $66.63 per hour, depending on experience, location, and employer.

What is a site reliability engineer?

A site reliability engineer specializes in site reliability engineering, or SRE, a specific branch of operations first pioneered by Google. You are responsible for ensuring that when a website decides to scale a particular feature for various users to access, it does not break the underlying software or website functions. This means you need to use analytical problem-solving skills to determine how to make specific features on a new software release work on top of existing source code.

What is a site reliability engineer?

A Site Reliability Engineer (SRE) is a professional who applies software engineering principles to infrastructure and operations problems. Their primary goal is to create scalable and highly reliable software systems, often bridging the gap between development and IT operations. SREs automate tasks, monitor system health, respond to incidents, and work to improve system reliability and performance. They also help define service level objectives (SLOs) and ensure systems meet customer expectations for uptime and availability.

What are the key skills and qualifications needed to thrive as a site reliability engineer?

To thrive as a Site Reliability Engineer, you need a strong background in computer science, systems administration, and software engineering, often supported by a degree in a technical field. Familiarity with cloud platforms (like AWS or GCP), container orchestration (such as Kubernetes), infrastructure as code (Terraform or Ansible), and monitoring tools (Prometheus, Grafana) is typically expected. Strong problem-solving skills, effective communication, and a proactive mindset help SREs excel at incident management and cross-functional collaboration. These skills are crucial for maintaining system reliability, minimizing downtime, and driving continuous improvement in complex technical environments.

What are some of the most common challenges site reliability engineers face when balancing system reliability with rapid software delivery?

Site Reliability Engineers (SREs) often navigate the challenge of maintaining highly reliable systems while supporting fast-paced software releases. This involves managing incidents, automating processes to reduce manual toil, and working closely with development teams to embed reliability into the software development lifecycle. SREs must carefully prioritize their efforts between proactive improvements and urgent, reactive fire-fighting. Effective communication and collaboration with both operations and development teams are crucial to ensuring service uptime without slowing down innovation.

What is the difference between Site Reliability Engineer vs DevOps Engineer?

AspectSite Reliability EngineerDevOps Engineer
CredentialsTypically requires a computer science degree, certifications like AWS, Google Cloud, or KubernetesSimilar credentials, often with cloud certifications and scripting skills
Work EnvironmentFocuses on maintaining and improving system reliability, often in large-scale production environmentsWorks on automation, CI/CD pipelines, and deployment processes across development and operations teams
Industry UsageCommon in tech, cloud services, and large-scale enterprise companiesWidely used in software development, cloud, and IT organizations

Both roles require strong technical skills and cloud knowledge, but SREs focus more on system reliability and uptime, while DevOps engineers emphasize automation and deployment processes. They often collaborate but have distinct primary responsibilities.

Is a site reliability engineer a stressful job?

A site reliability engineer (SRE) role can be stressful due to the responsibility of maintaining system uptime, handling incidents, and ensuring reliability under tight deadlines. The job often requires strong problem-solving skills, familiarity with monitoring tools, and the ability to work in high-pressure situations, but it also offers opportunities for skill development and process improvements.

What are the most commonly searched types of Site Reliability Engineer jobs in Detroit, MI?

The most popular types of Site Reliability Engineer jobs in Detroit, MI are:

What job categories do people searching Site Reliability Engineer jobs in Detroit, MI look for?

The top searched job categories for Site Reliability Engineer jobs in Detroit, MI are:

What cities near Detroit, MI are hiring for Site Reliability Engineer jobs?

Cities near Detroit, MI with the most Site Reliability Engineer job openings:

Infographic showing various Site Reliability Engineer job openings in Detroit, MI as of August 2026, with employment types broken down into 1% As Needed, 82% Full Time, 15% Part Time, and 2% Contract. Highlights an 93% Physical, 2% Hybrid, and 5% Remote job distribution, with an average salary of $121,312 per year, or $58.3 per hour.

SRE Engineer

Dearborn, MI โ€ข On-site

Ford Motor Company
Civil Engineering Constructionย โ€ขย 51 - 200 employees

$52.75 - $70/hr

Full-time

Re-posted 9 days ago


Job description

As a Site Reliability Engineer at Ford Motor Company, you will play a pivotal role in elevating the performance and dependability of our Marketing and Sales Tech platform and applications. In this essential position, your responsibilities will include closely collaborating with diverse teams across the organization to fortify our systems, ensuring they are not only robust and scalable but also equipped to efficiently manage the complexities of a global customer base. Your expertise in site reliability will be crucial in driving ongoing enhancements to our technology landscape. This continuous improvement effort is vital to maintaining Ford's leadership in innovation within the automotive industry, helping us set standards in service reliability and satisfaction. Your contributions will directly impact the smooth operation and evolutionary growth of our MS Tech capabilities, aligning with Ford's commitment to excellence and innovation.

  • Education: Bachelor's in computer science or related field.ย 
  • Experience: Minimum of 5+ years of professional experience in Site Reliability Engineering, or DevOps.ย 
  • Cloud Fluency (GCP): Deep hands-on experience with Google Cloud Platform, specifically Cloud Run, GKE, and OpenShift. You understand the nuances of container orchestration and serverless architectures.ย 
  • Infrastructure as Code (IaC): Advanced proficiency in Terraform. You must have experience writing reusable modules, managing state, and automating infrastructure provisioning (not just running existing scripts).ย 
  • Observability Engineering: Experience in comprehensive system observability using primary telemetry - Metrics, Events, Logs and Traces.ย ย 
    • Hands-on experience with Dynatrace (or similar APM tools like Datadog/New Relic), including distributed tracing, synthetic monitoring, and code-level profiling.
  • Coding & Scripting: Proficiency in at least one high-level programming language (Java, Node.js, Python, or Go). You will need to read application code to assist with instrumentation and debug complex production issues.ย 
  • Incident Management: Proven experience managing high-severity incidents. You understand the lifecycle of an incident: Triage -> Mitigation -> Resolution -> Blameless Post-Mortem (RCA).

Grade 7 or 8.

#LI-On-Site
#LI-DS2ย 

Incident Management & Operational Excellence

  • 24/7 Response: Participate in a 24/7 on-call rotation, providing rapid response to critical incidents and ensuring high availability for the NA eCommerce platform.
  • Incident Triage: Act as a decisive member of triage teams to diagnose, troubleshoot, and resolve complex production issues, directly contributing to the reduction of Mean Time to Recovery (MTTR).
  • Standardization: Diligently execute and contribute to the continuous improvement of operational runbooks and Standard Operating Procedures (SOPs) to ensure consistent incident response.

Problem Management & Root Cause Analysis

  • Blameless Culture: Lead and participate in blameless post-mortems and Root Cause Analysis (RCA) sessions to identify systemic weaknesses and implement preventative measures.
  • Strategic Engineering: Partner with cross-functional development and platform teams to architect long-term reliability solutions based on RCA findings.

Service Level Management (SLIs, SLOs, & Error Budgets)

  • Reliability Framework: Define and track meaningful Service Level Indicators (SLIs) and Objectives (SLOs) to measure service availability and performance.
  • Stakeholder Alignment: Collaborate with Product Owners to establish acceptable service levels and manage Error Budgets to balance velocity with reliability.
  • Performance Insight: Provide critical analysis during monthly release reviews, evaluating the impact of changes on service health and SLO adherence.

Modern Observability & Monitoring

  • Full-Stack Visibility: Leverage and optimize Ford's observability suite (Dynatrace, GCP Logging, etc.) to monitor system health and proactively identify anomalies.
  • Gap Remediation: Identify and document observability blind spots, implementing technical solutions to ensure comprehensive system visibility.
  • Monitoring-as-Code: Manage and refine metric collection, dashboard creation, and alert definitions, utilizing Terraform to provision monitoring infrastructure following Ford standards.
  • Alerting Strategy: Design robust notification strategies and thresholds to alert stakeholders of KPI/SLO violations using Error Budget signals.

Automation & Reliability Engineering

  • Toil Reduction: Champion the elimination of manual, repetitive tasks by developing automation scripts, tools, and streamlined workflows.
  • Resilience Engineering: Design and implement self-healing mechanisms to automatically detect and remediate common system failures, reducing the need for manual intervention.
  • Next-Gen Tech: Implement and manage AI-driven observability solutions to enhance proactive system monitoring and predictive maintenance.

Collaboration & Communication

  • Cross-Functional Leadership: Coordinate with platform and engineering teams to resolve production bottlenecks and drive continuous process improvements.
  • Strategic Reporting: Deliver clear, data-driven status reports on system health, incident trends, and SRE initiatives to leadership and program stakeholders.

Ford logo

About Ford

Sourced by ZipRecruiter

At Ford Motor Company, we believe freedom of movement drives human progress. With our incredible plans for the future of mobility, we have a wide variety of opportunities for you to accelerate your career and help us define tomorrow's transportation.

Industry

Civil engineering construction

Company size

51 - 200 Employees

Headquarters location

Doral, FL, US

Year founded

1982