Summary
Our client is a well-established, top-tier financial institution with a national footprint and a reputation for stability and innovation in the banking sector. They're building out their Site Reliability Engineering practice and looking for a Senior SRE to help shape it from the ground up โ someone who wants real influence over how reliability is done, not just another cog in an established machine.
Responsibilities
- Design and implement advanced monitoring frameworks, and drive adoption of monitoring best practices across engineering teams.
- Partner with application teams to improve observability and get ahead of incidents before they happen.
- Lead incident response for high-severity events, providing technical direction that keeps resolution fast and disruption minimal.
- Evangelize SRE principles and help build out the practice's technical standards as it matures.
- Automate manual monitoring and response tasks; champion infrastructure-as-code initiatives and drive toil reduction across the org.
- Own post-incident reviews and drive follow-through on corrective actions.
- Analyze performance trends and capacity utilization to surface improvement opportunities.
- Support 24/7 monitoring operations for mission-critical systems processing significant transaction volume.
- Optimize alerting strategies to cut noise and sharpen mean time to detection.
- Maintain and evolve monitoring dashboards and reporting.
- Evaluate high-complexity changes and establish risk mitigation strategies.
- Build out SLIs/SLOs for critical services in partnership with product teams, and introduce error budget concepts to guide reliability/velocity trade-off conversations.
- Foster knowledge sharing and document tribal knowledge across the team.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field โ or equivalent experience.
- 5+ years of experience in software development or site reliability engineering.
- Experience with high-availability, mission-critical systems.
- Proficiency with modern monitoring platforms.
- Strong troubleshooting and crisis-response capabilities.
- Broad development range โ comfortable from scripting through application development to infrastructure-as-code.
- Familiarity with DORA metrics and DevOps measurement frameworks.
- Solid understanding of network protocols, databases, and system architecture.
- Understanding of deployment pipeline reliability and release engineering practices.
- Strong communication skills for cross-team coordination.