Systems Engineer
Detroit, MI · Hybrid
Infrastructure Engineer (SRE) Department: Engineering - Infrastructure About Atomic Industries ... Design systems for secure, fault-tolerant deployment and observability * Build internal tooling to ...
Detroit, MI · Hybrid
Infrastructure Engineer (SRE) Department: Engineering - Infrastructure About Atomic Industries ... Design systems for secure, fault-tolerant deployment and observability * Build internal tooling to ...
Detroit, MI · Hybrid
Infrastructure Engineer (SRE) Department: Engineering - Infrastructure About Atomic Industries ... Design systems for secure, fault-tolerant deployment and observability * Build internal tooling to ...
Dearborn, MI · On-site
$48.50 - $66.50/hr
Champion Site Reliability (SRE): Design, build, evolve, and automate the production operations of our API Gateways, embedding SRE principles to ensure exceptional availability, performance, and ...
Quick apply
Dearborn, MI · On-site
$48.50 - $66.50/hr
Champion Site Reliability (SRE): Design, build, evolve, and automate the production operations of our API Gateways, embedding SRE principles to ensure exceptional availability, performance, and ...
Dearborn, MI · On-site
$48.50 - $66.50/hr
Bring a true SRE mindset to our platform. You will design self-healing architectures, establish robust SLOs/SLIs, and implement deep observability to detect and mitigate anomalies before they affect ...
Dearborn, MI · On-site
$48.50 - $66.50/hr
Bring a true SRE mindset to our platform. You will design self-healing architectures, establish robust SLOs/SLIs, and implement deep observability to detect and mitigate anomalies before they affect ...
Dearborn, MI · On-site
$51.25 - $68.50/hr
... alerting, and observability solutions using Prometheus, Grafana, Datadog, or Cloud ... SRE principles for system reliabilityDrive root cause analysis and continuous improvement ...
Dearborn, MI · On-site
$51.25 - $68.50/hr
... alerting, and observability solutions using Prometheus, Grafana, Datadog, or Cloud ... SRE principles for system reliabilityDrive root cause analysis and continuous improvement ...
Troy, MI · Remote
$114K - $172K/yr
Provide actionable escalation details to SRE, DevOps, and Security teams. * Maintain runbooks, SOPs, and troubleshooting guides. * Communicate effectively during incidents and followups. The ...
Troy, MI · Remote
$114K - $172K/yr
Provide actionable escalation details to SRE, DevOps, and Security teams. * Maintain runbooks, SOPs, and troubleshooting guides. * Communicate effectively during incidents and followups. The ...
Detroit, MI · On-site +1
$56.50 - $75/hr
The Staff Infrastructure Reliability Engineer is responsible for the technical leadership of Redfin ... on reliability, observability, scalability, and security. * You will support large scale / high ...
Detroit, MI · On-site +1
$56.50 - $75/hr
The Staff Infrastructure Reliability Engineer is responsible for the technical leadership of Redfin ... on reliability, observability, scalability, and security. * You will support large scale / high ...
... observability tools. * Collaboration and Leadership : Set up the necessary processes and ... reliability engineering (SRE). * Software Development: Experiencedeveloping, deploying ...
... observability tools. * Collaboration and Leadership : Set up the necessary processes and ... reliability engineering (SRE). * Software Development: Experiencedeveloping, deploying ...
Southfield, MI · Hybrid
$95K - $131K/yr
Experience with observability tools such as Prometheus, Grafana, and New Relic. * Experience with ... Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems ...
Quick apply
Southfield, MI · Hybrid
$95K - $131K/yr
Experience with observability tools such as Prometheus, Grafana, and New Relic. * Experience with ... Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems ...
Southfield, MI · Hybrid
$95K - $131K/yr
Experience with observability tools such as Prometheus, Grafana, and New Relic. * Experience with ... Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems ...
Southfield, MI · Hybrid
$95K - $131K/yr
Experience with observability tools such as Prometheus, Grafana, and New Relic. * Experience with ... Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems ...
Battle Creek, MI · On-site
$92K - $116K/yr
As a Reliability Engineer, you'll be the go-to partner for maintenance and operations, using data ... This role is designed for on-site partnership with maintenance and operations teams, with daily ...
Battle Creek, MI · On-site
$92K - $116K/yr
As a Reliability Engineer, you'll be the go-to partner for maintenance and operations, using data ... This role is designed for on-site partnership with maintenance and operations teams, with daily ...
Southfield, MI · Hybrid
$95K - $131K/yr
Experience with observability tools such as Prometheus, Grafana, and New Relic. * Experience with ... Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems ...
Southfield, MI · Hybrid
$95K - $131K/yr
Experience with observability tools such as Prometheus, Grafana, and New Relic. * Experience with ... Site Reliability Engineering (SRE) experience. * Experience with Active Directory and file systems ...
... SRE teams. How we work * Build small, ship often: incremental changes with clear acceptance criteria and fast feedback. * Quality is non-negotiable: tests, clear design, observability, and secure ...
Quick apply
... SRE teams. How we work * Build small, ship often: incremental changes with clear acceptance criteria and fast feedback. * Quality is non-negotiable: tests, clear design, observability, and secure ...
$40.86 - $57.70/hr
Reliability Engineer The Reliability Engineer will be responsible for identifying and managing ... Provide documentation and on-site training for maintenance and engineering personnel. Essential ...
Quick apply
$40.86 - $57.70/hr
Reliability Engineer The Reliability Engineer will be responsible for identifying and managing ... Provide documentation and on-site training for maintenance and engineering personnel. Essential ...
Detroit, MI · On-site
... observability (Better Stack, Prometheus, VictoriaMetrics, Grafana) · Assist with incident response and participate in an on-call rotation as the team grows · Work with SRE to improve system ...
Quick apply
Detroit, MI · On-site
... observability (Better Stack, Prometheus, VictoriaMetrics, Grafana) · Assist with incident response and participate in an on-call rotation as the team grows · Work with SRE to improve system ...
$112K - $148K/yr
Observability & Monitoring Build and maintain monitoring, alerting, and dashboarding solutions ... Operational Support Participate in on-call rotations, troubleshoot incidents, and apply SRE ...
Quick apply
$112K - $148K/yr
Observability & Monitoring Build and maintain monitoring, alerting, and dashboarding solutions ... Operational Support Participate in on-call rotations, troubleshoot incidents, and apply SRE ...
Dearborn, MI · On-site
Experience with Kubernetes and infrastructure as code in partnership with platform/SRE teams.Quality is non-negotiable: tests, clear design, observability, and secure defaults. Experience Required6 ...
Dearborn, MI · On-site
Experience with Kubernetes and infrastructure as code in partnership with platform/SRE teams.Quality is non-negotiable: tests, clear design, observability, and secure defaults. Experience Required6 ...
Ann Arbor, MI · On-site
$102K - $140K/yr
Reliability: Define SLOs, improve observability, and lead incident response and root cause analysis ... Qualifications * 5+ years in Platform Engineering / SRE / DevOps * Strong Azure expertise (AKS ...
Ann Arbor, MI · On-site
$102K - $140K/yr
Reliability: Define SLOs, improve observability, and lead incident response and root cause analysis ... Qualifications * 5+ years in Platform Engineering / SRE / DevOps * Strong Azure expertise (AKS ...
Dearborn, MI · On-site +1
$48.50 - $66.50/hr
Bring a true SRE mindset to our platform. You will design self-healing architectures, establish robust SLOs/SLIs, and implement deep observability to detect and mitigate anomalies before they affect ...
Quick apply
Dearborn, MI · On-site +1
$48.50 - $66.50/hr
Bring a true SRE mindset to our platform. You will design self-healing architectures, establish robust SLOs/SLIs, and implement deep observability to detect and mitigate anomalies before they affect ...
$102K - $140K/yr
Reliability: Define SLOs, improve observability, and lead incident response and root cause analysis ... Qualifications * 5+ years in Platform Engineering / SRE / DevOps * Strong Azure expertise (AKS ...
Quick apply
$102K - $140K/yr
Reliability: Define SLOs, improve observability, and lead incident response and root cause analysis ... Qualifications * 5+ years in Platform Engineering / SRE / DevOps * Strong Azure expertise (AKS ...
Dearborn, MI · On-site
$99K - $166K/yr
Exposure to SRE practices and reliability engineering. * Experience supporting ML or data-intensive ... Improve reliability and observability of CI/CD systems through monitoring, logging, and alerting.
Dearborn, MI · On-site
$99K - $166K/yr
Exposure to SRE practices and reliability engineering. * Experience supporting ML or data-intensive ... Improve reliability and observability of CI/CD systems through monitoring, logging, and alerting.
| Aspect | Observability Site Reliability Engineer | Monitoring Engineer |
|---|---|---|
| Focus | Ensuring system reliability through observability, automation, and incident response | Implementing and managing monitoring tools and dashboards |
| Skills | Cloud platforms, scripting, incident management, observability tools | Monitoring tools, alerting systems, data analysis |
| Work Environment | DevOps teams, cloud infrastructure, large-scale systems | Operations teams, infrastructure monitoring |
While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.
Other
Medical, Dental, Vision, Retirement, PTO
Re-posted 22 days ago
Atomic Industries is reinventing how the world makes things. From cars and aerospace systems to medical devices and packaging, most physical goods begin life in a mold or are shaped by a manufacturing tool. Producing these tools has always been slow, manual, and dependent on scarce expertise, taking weeks or months. We're changing that.
At our Detroit headquarters, we combine the industrial DNA of America's manufacturing heartland with the speed, intelligence, and precision of Silicon Valley. Our AI-driven platform tackles the hardest problems in geometry, process planning, and fabrication, collapsing production timelines from months to days and soon, minutes. We don't just build software; we run a fully operational factory where our technology produces production-grade tooling every week, enabling tight feedback loops and rapid iteration.
Backed by top-tier investors, we're restoring speed, flexibility, and capability to the American industrial base. Our mission is to make manufacturing as agile and scalable as the digital world, and in doing so, rebuild the infrastructure of the physical economy.
As an Infrastructure Engineer at Atomic, you'll own the systems that power high-performance simulation, geometry processing, and factory automation. You'll be responsible for building and maintaining the hybrid infrastructure that spans cloud, on-prem, and GPU compute environments.
This role is ideal for engineers who thrive at the intersection of reliability, performance, and engineering velocity.
Manage and improve CI/CD pipelines (GitHub Actions, Terraform, HashiCorp stack)
Orchestrate compute workloads across cloud and bare-metal GPU clusters
Design systems for secure, fault-tolerant deployment and observability
Build internal tooling to accelerate developer workflows and operations
Partner with engineers on simulation, ML, and backend teams to scale infrastructure
Monitor and tune system performance under high computational loads
5+ years in DevOps, infrastructure, or site reliability engineering
Proficiency with container orchestration (Nomad, Kubernetes), IaaC, and Linux
Strong scripting and programming skills (Python, Bash, Go)
Experience with hybrid (cloud + on-prem) infrastructure
Familiarity with distributed systems, caching layers, and observability tools
Background supporting ML, simulation, or HPC workloads
Experience managing large GPU fleets or multi-node compute tasks
Security or compliance experience (ITAR, SOC2, CMMC)
Track record of building tooling for high-autonomy engineering teams
Fast iteration: We deploy code daily and validate it in the factory every week
Factory-first mindset: Engineers spend time on the shop floor to understand the real-world impact of their work
Low ceremony, high ownership: We value well-written design docs, tested code, and real results - not meetings
Collaborative culture: Geometry, AI, and robotics teams are tightly integrated and co-develop the product together
Competitive salary and generous equity package
Full medical, dental, and vision coverage for employees and dependents
401(k)
PTO with a 15-day minimum
Quarterly team travel to Detroit
Visa support for TN, E3, and O1 candidates
Hardware stipend and on-site prototype lab access
Sourced by ZipRecruiter
Internet and it
11 - 50 Employees
Dripping Springs, TX, US