Overview
Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the reliability, scalability, and performance of critical Infrastructure Automation platforms. This role will lead the design and implementation of SRE practices across a federated technology ecosystem, ensuring operational excellence through automation, observability, and resilient architecture.
The ideal candidate will bring deep expertise in distributed systems, cloud‑native infrastructure, SaaS application support and DevOps/SRE principles, along with strong leadership and collaboration skills to influence cross‑functional engineering and Production management teams and drive continuous improvement in service reliability.
Responsibilities
SRE Strategy & Governance:
- Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols.
- Establish governance models for reliability engineering across distributed teams.
- Champion a culture of observability, proactive monitoring, and continuous feedback loops.
Reactive & Proactive Problem Management:
- Lead root cause analysis (RCA) and post‑incident reviews to identify systemic issues and prevent recurrence.
- Implement proactive problem detection using telemetry, anomaly detection, and trend analysis.
- Collaborate with engineering and operations teams to eliminate toil and reduce incident frequency and impact.
Capacity & Performance Management:
- Develop and maintain capacity models to ensure systems scale efficiently with business demand.
- Monitor performance trends and lead optimization efforts across infrastructure and applications.
- Partner with finance and engineering teams to align capacity planning with cost and growth objectives.
Platform Reliability & Automation:
- Drive automation of operational tasks including deployments, scaling, and recovery.
- Integrate reliability tooling with CI/CD pipelines, ITSM platforms (ServiceNow), and observability systems.
Incident Management & Operational Excellence:
- Oversee major incident response, escalation, and communication processes.
- Develop and maintain runbooks, playbooks, and escalation protocols.
- Drive continuous improvement through blameless retrospectives and operational reviews.
Technical Leadership:
- Serve as a senior technical advisor and thought leader in SRE and platform engineering.
- Mentor and guide SRE teams and partner with engineering leaders across the enterprise.
- Provide input on staffing, tooling strategy, and budget planning for reliability initiatives.
Managerial Responsibilities
- Opportunity & Inclusion Champion: Models an inclusive environment for employees and clients, aligned to company Great Place to Work goals.
- Manager of Process & Data: Demonstrates deep process knowledge, operational excellence and innovation through a focus on simplicity, data based decision making and continuous improvement.
- Enterprise Advocate & Communicator: Communicates enterprise decisions, purpose, and results, and connects to team strategy, priorities and contributions.
- Risk Manager: Ensures proper risk discipline, controls and culture are in place to identify, elevate and debate issues.
- People Manager & Coach: Provides inspection, coaching and feedback to motivate, differentiate and improve performance.
- Financial Steward: Actively manages expenses and budgets in alignment with objectives, making sound financial decisions.
- Enterprise Talent Leader: Assesses talent and builds bench strength for roles across the organization.
- Driver of Business Outcomes: Delivers results by effectively prioritizing, inspecting and appropriately delegating team work.
Required Qualifications
- 10+ years of experience in systems engineering, DevOps, or SRE roles in large‑scale environments.
- Deep understanding of Linux/Unix & Windows systems, networking, and distributed computing.
- Proven experience with observability stacks (Dynatrace, Grafana, Splunk, OpenTelemetry).
- Expertise in infrastructure‑as‑code and automation tools (Terraform, Ansible, Python).
- Strong knowledge of cloud platforms and container orchestration (Kubernetes).
- Demonstrated success in leading incident response and driving systemic improvements.
- Experience with capacity planning, performance tuning, and cost optimization.
- Excellent communication and stakeholder management skills, including executive engagement.
Desired Qualifications
- Experience with ITIL/ITSM processes and integration with platforms like ServiceNow.
- Familiarity with security and compliance in regulated industries (financial services).
- Background in performance engineering and infrastructure analytics.
- Experience developing dashboards and metrics for operational health and reliability.
Skills
- Influence
- Risk Management
- Solution Design
- Stakeholder Management
- Technical Strategy Development
- Analytical Thinking
- Application Development
- Collaboration
- Result Orientation
- Solution Delivery Process
- Agile Practices
- Architecture
- Automation
- Data Management
- DevOps Practices
Shift
1st shift (United States of America)
Hours Per Week
40
#J-18808-Ljbffr