Must Haves:
4+ years experience SRE
2+ years experience with monitoring and observability platforms (NewRelic, Prometheus, Grafana, etc.)
2+ years Cloud Experience, Chaos Engineering experience preferred
Job Description:
The InfoSec SRE is a pivotal role focused on engineering and advancing the maturity of the organization''s site reliability framework development, adoption, and integration. This position establishes the foundational framework for Site Reliability Engineering (SRE) design fabrics within InfoSec. Additionally, the role supports management of runbook repositories, observability, and related technologies, acting as a trusted advisor to peers and stakeholders across the organization.
Essential/Key Responsibilities:
โข Design and implement foundational SRE practices (SLIs/SLOs, error budgets, incident management)
โข Partner with InfoSec and engineering teams to define reliability standards and operating models
โข Establish and drive adoption of SRE principles across teams and provide guidance to onboarding organizations to the framework.
โข Design and implement incident response processes and escalation models
โข Integrate and optimize alerting and on-call workflows using PagerDuty
โข Develop and maintain operational runbooks and playbooks
โข Lead or support incident reviews and postmortems with a focus on continuous improvement
โข Integrate SRE workflows with enterprise platforms such as ServiceNow
โข Automate operational tasks, incident workflows, and reporting
โข Improve system resilience through automation and self-healing mechanisms
โข Design and execute chaos engineering experiments to validate system resilience
โข Identify failure modes and proactively address system weaknesses
โข Collaborate with engineering teams to improve fault tolerance and recovery strategies
โข Drive adoption of reliability best practices across the organization
โข Provide guidance and mentorship on SRE principles
Required Skills/Knowledge
โข 4-5 years of experience with SRE, DevOps, or Infrastructure engineering
โข Skills with multi-cloud environments (AWS, Azure) as well as on-prem integrations
โข Strong experience with monitoring and observability platforms (NewRelic, Prometheus, Grafana, etc.)
โข Experience with incident management and on-call systems (e.g. PagerDuty)
โข Familiarity with ITSM platforms, specifically ServiceNow and integration via API
โข Solid understanding of system reliability, performance, and scalability
โข Automation scripting
โข Runbook generation
Desired Skills/Knowledge:
โข Experience working within or alongside InfoSec teams
โข Proficiency with scripting languages such as Python, Bash, Ruby etc.
โข Knowledge of security or security-adjacent tools, controls, and compliance frameworks
โข Experience implementing SLO''s, SLAs, and error budgeting
โข Exposure to chaos engineering practices (specifically, SteadyBit as a tool)
โข Strong communication and collaboration skills, specifically documenting
โข Systems thinking and problem-solving mindset
โข A strong focus on automation and continuous improvement methodologies
โข Data-driven decision making
โข Ability to operate in ambiguous, greenfield environments
โข Understanding of Infrastructure as Code
โข Passion for the work and responsibility to the consumer
โข Understanding of public cloud platforms and services
โข Positive attitude with a strong desire to continuously learn and adapt