1

Observability Site Reliability Engineer Jobs in Michigan

Infrastructure Engineer (SRE) Department: Engineering - Infrastructure About Atomic Industries ... Design systems for secure, fault-tolerant deployment and observability * Build internal tooling to ...

DevOps Engineer

Dearborn, MI · On-site

$48.50 - $66.50/hr

Bring a true SRE mindset to our platform. You will design self-healing architectures, establish robust SLOs/SLIs, and implement deep observability to detect and mitigate anomalies before they affect ...

Cloud Engineer

Dearborn, MI · On-site

$51.25 - $68.50/hr

... alerting, and observability solutions using Prometheus, Grafana, Datadog, or Cloud ... SRE principles for system reliabilityDrive root cause analysis and continuous improvement ...

Senior Systems Engineer

Southfield, MI · Hybrid

$95K - $131K/yr

Experience with observability tools such as Prometheus, Grafana, and New Relic. * Experience with ... Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems ...

Senior Systems Engineer

Southfield, MI · Hybrid

$95K - $131K/yr

Experience with observability tools such as Prometheus, Grafana, and New Relic. * Experience with ... Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems ...

Reliability Engineer

Battle Creek, MI · On-site

$92K - $116K/yr

As a Reliability Engineer, you'll be the go-to partner for maintenance and operations, using data ... This role is designed for on-site partnership with maintenance and operations teams, with daily ...

Senior Systems Engineer

Southfield, MI · Hybrid

$95K - $131K/yr

Experience with observability tools such as Prometheus, Grafana, and New Relic. * Experience with ... Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems ...

... SRE teams. How we work * Build small, ship often: incremental changes with clear acceptance criteria and fast feedback. * Quality is non-negotiable: tests, clear design, observability, and secure ...

Reliability Engineer The Reliability Engineer will be responsible for identifying and managing ... Provide documentation and on-site training for maintenance and engineering personnel. Essential ...

... observability (Better Stack, Prometheus, VictoriaMetrics, Grafana) · Assist with incident response and participate in an on-call rotation as the team grows · Work with SRE to improve system ...

Experience with Kubernetes and infrastructure as code in partnership with platform/SRE teams.Quality is non-negotiable: tests, clear design, observability, and secure defaults. Experience Required6 ...

Senior Platform Engineer

Ann Arbor, MI · On-site

$102K - $140K/yr

Reliability: Define SLOs, improve observability, and lead incident response and root cause analysis ... Qualifications * 5+ years in Platform Engineering / SRE / DevOps * Strong Azure expertise (AKS ...

DevOps Specialist

Dearborn, MI · On-site +1

$48.50 - $66.50/hr

Bring a true SRE mindset to our platform. You will design self-healing architectures, establish robust SLOs/SLIs, and implement deep observability to detect and mitigate anomalies before they affect ...

Senior Platform Engineer

Ann Arbor, MI

$102K - $140K/yr

Reliability: Define SLOs, improve observability, and lead incident response and root cause analysis ... Qualifications * 5+ years in Platform Engineering / SRE / DevOps * Strong Azure expertise (AKS ...

Showing results 21-40

Observability Site Reliability Engineer information

What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?

AspectObservability Site Reliability EngineerMonitoring Engineer
FocusEnsuring system reliability through observability, automation, and incident responseImplementing and managing monitoring tools and dashboards
SkillsCloud platforms, scripting, incident management, observability toolsMonitoring tools, alerting systems, data analysis
Work EnvironmentDevOps teams, cloud infrastructure, large-scale systemsOperations teams, infrastructure monitoring

While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.

What job categories do people searching Observability Site Reliability Engineer jobs in Michigan look for? The top searched job categories for Observability Site Reliability Engineer jobs in Michigan are:
What cities in Michigan are hiring for Observability Site Reliability Engineer jobs? Cities in Michigan with the most Observability Site Reliability Engineer job openings:

Other

Medical, Dental, Vision, Retirement, PTO

Re-posted 22 days ago


Job description

Infrastructure Engineer (SRE)Department: Engineering - Infrastructure
About Atomic Industries

Atomic Industries is reinventing how the world makes things. From cars and aerospace systems to medical devices and packaging, most physical goods begin life in a mold or are shaped by a manufacturing tool. Producing these tools has always been slow, manual, and dependent on scarce expertise, taking weeks or months. We're changing that.

At our Detroit headquarters, we combine the industrial DNA of America's manufacturing heartland with the speed, intelligence, and precision of Silicon Valley. Our AI-driven platform tackles the hardest problems in geometry, process planning, and fabrication, collapsing production timelines from months to days and soon, minutes. We don't just build software; we run a fully operational factory where our technology produces production-grade tooling every week, enabling tight feedback loops and rapid iteration.

Backed by top-tier investors, we're restoring speed, flexibility, and capability to the American industrial base. Our mission is to make manufacturing as agile and scalable as the digital world, and in doing so, rebuild the infrastructure of the physical economy.


About the Role

As an Infrastructure Engineer at Atomic, you'll own the systems that power high-performance simulation, geometry processing, and factory automation. You'll be responsible for building and maintaining the hybrid infrastructure that spans cloud, on-prem, and GPU compute environments.

This role is ideal for engineers who thrive at the intersection of reliability, performance, and engineering velocity.


What You'll Do
  • Manage and improve CI/CD pipelines (GitHub Actions, Terraform, HashiCorp stack)

  • Orchestrate compute workloads across cloud and bare-metal GPU clusters

  • Design systems for secure, fault-tolerant deployment and observability

  • Build internal tooling to accelerate developer workflows and operations

  • Partner with engineers on simulation, ML, and backend teams to scale infrastructure

  • Monitor and tune system performance under high computational loads


What We're Looking ForMinimum Qualifications
  • 5+ years in DevOps, infrastructure, or site reliability engineering

  • Proficiency with container orchestration (Nomad, Kubernetes), IaaC, and Linux

  • Strong scripting and programming skills (Python, Bash, Go)

  • Experience with hybrid (cloud + on-prem) infrastructure

  • Familiarity with distributed systems, caching layers, and observability tools

Bonus Points
  • Background supporting ML, simulation, or HPC workloads

  • Experience managing large GPU fleets or multi-node compute tasks

  • Security or compliance experience (ITAR, SOC2, CMMC)

  • Track record of building tooling for high-autonomy engineering teams


How We Work
  • Fast iteration: We deploy code daily and validate it in the factory every week

  • Factory-first mindset: Engineers spend time on the shop floor to understand the real-world impact of their work

  • Low ceremony, high ownership: We value well-written design docs, tested code, and real results - not meetings

  • Collaborative culture: Geometry, AI, and robotics teams are tightly integrated and co-develop the product together


Benefits
  • Competitive salary and generous equity package

  • Full medical, dental, and vision coverage for employees and dependents

  • 401(k)

  • PTO with a 15-day minimum

  • Quarterly team travel to Detroit

  • Visa support for TN, E3, and O1 candidates

  • Hardware stipend and on-site prototype lab access