1

Observability Site Reliability Engineer Jobs in Missouri

Core Engineering: 10+ years of experience in SRE, DevOps, or Software/Systems Engineering ... Observability: Hands-on experience with monitoring, logging, and tracing stacks (e.g., Prometheus ...

Lead Site Reliability Engineer

O Fallon, MO · On-site

$53.25 - $70.75/hr

You see CI/CD, automation, observability, and reliability as foundational engineering disciplines ... on-site fitness facilities; eligibility for tuition reimbursement; and many more. Mastercard ...

Lead Site Reliability Engineer

O Fallon, MO · On-site

$53.25 - $70.75/hr

You see CI/CD, automation, observability, and reliability as foundational engineering disciplines ... on-site fitness facilities; eligibility for tuition reimbursement; and many more. Mastercard ...

Lead Site Reliability Engineer

O Fallon, MO · On-site

$53.25 - $70.75/hr

You see CI/CD, automation, observability, and reliability as foundational engineering disciplines ... on-site fitness facilities; eligibility for tuition reimbursement; and many more. Mastercard ...

Lead Site Reliability Engineer

O Fallon, MO · On-site

$53.25 - $70.75/hr

You see CI/CD, automation, observability, and reliability as foundational engineering disciplines ... on-site fitness facilities; eligibility for tuition reimbursement; and many more. Mastercard ...

Senior Site Reliability Engineer

Berkeley, MO · On-site

$53.50 - $71/hr

Senior Site Reliability Engineer Company: The Boeing Company The Boeing Company is looking for a Senior Site Reliability Engineer to join the Air Dominance Site Reliability Engineering team located ...

Showing results 21-40

Observability Site Reliability Engineer information

What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?

AspectObservability Site Reliability EngineerMonitoring Engineer
FocusEnsuring system reliability through observability, automation, and incident responseImplementing and managing monitoring tools and dashboards
SkillsCloud platforms, scripting, incident management, observability toolsMonitoring tools, alerting systems, data analysis
Work EnvironmentDevOps teams, cloud infrastructure, large-scale systemsOperations teams, infrastructure monitoring

While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.

What job categories do people searching Observability Site Reliability Engineer jobs in Missouri look for? The top searched job categories for Observability Site Reliability Engineer jobs in Missouri are:
What cities in Missouri are hiring for Observability Site Reliability Engineer jobs? Cities in Missouri with the most Observability Site Reliability Engineer job openings:

Staff Site Reliability Engineer

Jobtailor

California, MO • On-site

$140 - $210/hr

Other

Posted 4 days ago


Job description

  • System Resilience: Design, build, and maintain highly available, scalable, and secure infrastructure to support our AI-native cybersecurity platform.
  • Automation & Tooling: Develop internal tooling and automation to streamline deployment processes, incident response, and capacity planning.
  • Performance Engineering: Monitor system performance and proactively identify bottlenecks, optimizing infrastructure for low-latency, high-throughput AI workloads.
  • Incident Management: Lead incident response efforts, conduct post-mortems, and implement long-term solutions to prevent recurring reliability issues.
  • Infrastructure as Code (IaC): Manage infrastructure via code, driving consistency, auditability, and scalability across our cloud environments (e.g., AWS, GCP).
  • Cross-Functional Collaboration: Partner with sibling Engineering teams, Product, and Security teams to ensure reliability is baked into our development lifecycle from concept to production.
Requirements
  • Core Engineering: 10+ years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing production systems at scale.
  • Cloud Infrastructure: Deep expertise in public cloud environments (AWS, GCP, or Azure) and managing services such as Kubernetes (EKS/GKE), networking, and storage.
  • Infrastructure as Code: Extensive experience with tools like Terraform, Pulumi, or similar technologies to manage complex infrastructure deployments.
  • Observability: Hands-on experience with monitoring, logging, and tracing stacks (e.g., Prometheus, Grafana, ELK, Datadog) to drive data-informed reliability decisions.
  • Distributed Systems: Solid understanding of microservices architecture, distributed databases, and event-driven systems.
  • Communication: Clear, concise communication skills and a bias for collaborative problem-solving.
  • Leadership Alignment: Proven track record of guiding multi-stakeholder initiatives and influencing engineering practices across teams.
  • Analytical Rigor: Strong problem-solving, debugging, and analytical skills, especially in high-pressure environments.
  • Domain Background: Prior work in cybersecurity, specifically regarding SIEM, EDR, or SOAR infrastructure is nice-to-have.
  • AI/ML Infrastructure: Experience supporting infrastructure for large-scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization) is nice-to-have.
  • Startup Mentality: Background driving high-impact engineering initiatives in high-growth startups or enterprise SaaS is nice-to-have.
  • Strong familiarity with Agentic Workflows such as Agno, Temporal, etc. is nice-to-have.
Core Competencies

Demonstrates expertise in designing and maintaining scalable, secure cloud infrastructure, with a strong focus on automation, performance engineering, and incident management. Proven ability to collaborate across teams and drive reliability in AI-native cybersecurity platforms.

Highest-signal resume keywords
  • 10+ Years Experience in SRE, DevOps, or Software Engineering
  • Deep Expertise in AWS, GCP, or Azure
  • Extensive Experience with Terraform or Pulumi
  • Hands-On Experience with Prometheus, Grafana, or ELK
  • Strong Problem-Solving and Analytical Skills
ATS Optimization Keywords Hard Skills
  • Infrastructure as Code
  • Performance Engineering
  • Incident Management
  • Cloud Infrastructure Management
  • Distributed Systems Understanding
  • AI/ML Infrastructure Support
  • Automation Development
  • Monitoring and Observability
  • Capacity Planning
  • Microservices Architecture
Soft Skills
  • Clear Communication Skills
  • Collaborative Problem-Solving
  • Leadership in Multi-Stakeholder Initiatives
Industry Keywords
  • Cybersecurity
  • AI-Native Platforms
  • High-Impact Engineering
  • Startup Mentality
  • Enterprise SaaS
Tools & Technologies
  • Kubernetes (EKS/GKE)
  • Prometheus
  • Grafana
  • ELK
  • Datadog
  • Terraform
  • Pulumi
  • Agentic Workflows
  • SIEM
  • EDR
#J-18808-Ljbffr