Job Title: Senior Observability & Incident Response Standards Lead
Role Overview
We are seeking a Senior Observability & Incident Response Standards Lead to own the health, clarity, and efficacy of our monitoring and alerting ecosystem. In this role, you will be the final authority on what constitutes an "actionable" alert. You will define the standards for how our infrastructure notifies our response teams, ensuring that every alert is prioritized, well-documented, and tied to clear remediation instructions.
This is a governance and architecture role. You will sit at the intersection of Engineering, Infrastructure, and Operations, serving as the bridge between technical system health and business-level service reliability. If you are passionate about eliminating "alert fatigue," designing robust incident response frameworks, and building a culture of observability, we want to speak with you.
Key Responsibilities
Alert Rationalization & Signal Governance
• Framework Development: Build and maintain an enterprise-wide alert rationalization framework. Evaluate alerts based on criticality, actionability, and signal-to-noise ratio.
• Gatekeeping: Act as the primary governance lead for determining alerting thresholds. Ensure that notifications are routed correctly-whether to automated ticketing systems, business-hours support, or mission-critical response teams.
• Continuous Improvement: Lead recurring reviews to audit alert quality, suppress low-value noise, and ensure that every signal reaching a human is worth their time.
Standards, Policies & Knowledge Management
• Standardization: Define and enforce technical standards for severity, naming conventions, tagging taxonomy, and alert metadata (e.g., CI, service owner, runbook link).
• Response Cataloging: Establish a scalable approach to maintaining a "Knowledge System." Ensure every actionable alert is paired with a versioned, high-quality runbook that clearly outlines triage steps, impact, and escalation triggers.
• Operational Guardrails: Partner with tool/platform owners to embed these standards directly into our observability tooling (templates, required fields, and automated validation).
Strategy & Outcomes
• KPI Management: Define and publish metrics that demonstrate the health of our monitoring ecosystem (e.g., alert actionability rate, volume trends, runbook coverage).
• Enablement: Coach engineering teams on SRE principles, SLI/SLO alignment, and multi-signal alerting patterns.
• Feedback Loops: Use post-incident learnings to drive the evolution of our monitoring standards, ensuring our alerting logic constantly matures with our system architecture.
Qualifications & Requirements
Minimum Qualifications
• Experience Baseline: 5+ years of professional experience in IT Operations, SRE, Observability, or Incident Management.
• Technical Mastery: Proven success reducing alert noise and improving response quality in large-scale enterprise environments.
• Tooling Proficiency: Deep familiarity with modern observability and monitoring stacks (e.g., Datadog, Dynatrace, Splunk, Prometheus/Grafana, or similar).
• Governance Experience: Strong understanding of service ownership models, CMDB dependency mapping, and incident response lifecycles.
Preferred Attributes
• Experience designing or operating a centralized Operations Command Center or NOC model.
• Expertise in event correlation, deduplication, and automated noise-suppression workflows.
• Strong background in ITIL Event Management or SRE "Reliability" methodologies.
• Exceptional stakeholder management skills-you must be comfortable influencing engineering leads and service owners to adopt your standards.
Equal Opportunity Employer / Disabled / Protected Veterans
The Know Your Rights poster is available here:
https://www.eeoc.gov/sites/default/files/2023-06/22-088_EEOC_KnowYourRights6.12.pdf
The pay transparency policy is available here:
https://www.dol.gov/sites/dolgov/files/ofccp/pdf/pay-transp_%20English_formattedESQA508c.pdf
For temporary assignments lasting 13 weeks or longer, AllSTEM Connections is pleased to offer major medical, dental, vision, 401k and any statutory sick pay where required.
We are committed to working with and providing reasonable accommodations to individuals with disabilities. If you need a reasonable accommodation for any part of the employment process, please contact your staffing representative who will reach out to our HR team.
AllSTEM Connections participates in the E-Verify program in certain locations as required by law. Learn more about the E-Verify program.
https://e-verify.uscis.gov/web/media/resourcesContents/E-Verify_Participation_Poster_ES.pdf
We also consider for employment qualified applicants regardless of criminal histories, consistent with legal requirements, including, if applicable, the City of Los Angeles' Fair Chance Initiative for Hiring Ordinance. Pursuant to applicable state and municipal Fair Chance Laws and Ordinances, we will consider for employment-qualified applicants with arrest and conviction records, including, if applicable, the San Francisco Fair Chance Ordinance. For Los Angeles, CA applicants: Qualified applications with arrest or conviction records will be considered for employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act.
Additional Skills
(none specified)
AllSTEM Representative Contact Info
Account Executive:
Nichols
Branch Phone:
(909) 244-1777
Location:
Ontario, CA