Overview
The Site Reliability Engineer (SRE) โ Release amp; Operations is a hybrid technical and process-oriented role responsible for bridging the gap between software development and stable IT operations. This role focuses on owning and optimizing the release lifecycle, driving governance through the Change Advisory Board (CAB), and facilitating cross-team deployment communication.
In addition to release engineering, this position requires a hands-on technical expert who will develop automated tooling to reduce operational toil, actively participate in production support and on-call rotations, and ensure the high availability, security, and scalability of AWS-based cloud infrastructure in a compliant corporate environment.
Responsibilities
Release Management amp; Governance (CAB)
- Release Lifecycle Ownership: Lead and coordinate the end-to-end release lifecycle, including planning, scheduling, staging, deploying, and post-release validation.
- Change Advisory Board (CAB): Act as a key representative and technical coordinator on the Change Advisory Board, defending upcoming releases, evaluating architectural risk, and ensuring all compliance requirements are met prior to production deployment.
- Cross-Functional Communication: Serve as the primary point of contact for developers, QA, product management, and business stakeholders regarding deployment windows, release status, risk assessments, and rollback plans.
- Process Optimization: Standardize and mature release processes, transitioning manual gatekeeping into automated CI/CD guardrails and repeatable workflows.
Site Reliability amp; Production Support
- Production Support amp; On-Call: Provide hands-on tier-2 production support, ensuring operational stability and participating in the active engineering on-call rotation.
- Tooling amp; Automation Development: Design, build, and maintain internal scripts, custom tooling, and automated pipelines (Python, Bash, PowerShell) to reduce operational "toil" and streamline release operations.
- Incident Response amp; RCA: Respond promptly to system outages, service interruptions, and security alerts. Lead post-incident Root Cause Analysis (RCA) efforts and implement permanent preventive engineering solutions.
- Infrastructure-as-Code (IaC): Maintain, deploy, and scale AWS cloud infrastructure using Terraform, AWS CloudFormation, or Ansible to ensure environment parity and drift-free deployments.
- Observability amp; Monitoring: Configure, tune, and optimize monitoring and alerting systems (e.g., CloudWatch, Prometheus, Grafana) to provide comprehensive visibility into release health and production performance.
Security, Compliance, amp; Collaboration
- Compliance Alignment: Ensure all deployment and release activities strictly adhere to PCI, SOC2, and internal corporate security/governance standards.
- Collaboration: Work closely with software engineering, QA, and platform infrastructure teams to build "paved paths" for developers to ship software safely and rapidly.
- Documentation: Maintain pristine, audit-ready documentation for release logs, standard operating procedures (SOPs), runbooks, and incident timelines.
Essential Skills and Experience
- Experience: 3โ5 years of experience in SRE, DevOps, system administration, or Release/Operations engineering roles.
- Release Experience: Proven experience coordinating software releases, managing multi-tier deployment pipelines, and working within formal ITIL/Change Management frameworks (including active CAB participation).
- Education: Bachelorโs degree in Computer Science, Engineering, Information Technology, or a related technical field (or equivalent practical experience).
- AWS Expertise: Hands-on experience configuring, deploying, and maintaining AWS infrastructure and serverless architectures (EC2, S3, RDS, IAM, Lambda, API Gateway).
- Infrastructure-as-Code (IaC): Solid understanding and working experience with IaC tools such as Terraform or AWS CloudFormation.
- Scripting amp; Automation: Strong proficiency in scripting languages (especially Python and Bash) to write automation tooling and integrate systems.
- Incident Management: Demonstrated experience leading Root Cause Analysis (RCA) and participating in production on-call rotations.
- Communication: Exceptional written and verbal communication skills, with a proven ability to coordinate across highly technical development teams and business-oriented leaders.
Preferred Skills and Experience
- Certifications: AWS Certified SysOps Administrator, AWS Certified DevOps Engineer, or ITIL Foundation certifications.
- CI/CD Pipeline Tooling: Experience configuring and maintaining CI/CD systems (such as GitLab CI, GitHub Actions, Jenkins, or AWS CodePipeline).
- Observability Tools: Hands-on experience with modern monitoring, APM, and alerting tools (e.g., Datadog, Prometheus, Grafana, PagerDuty).
- Containerization: Experience with Docker and container orchestration platforms (Kubernetes, AWS ECS/EKS).
- Compliance Audits: Hands-on experience assisting with SOC2 or PCI-DSS audits, especially documenting evidence for code changes and release gates.
Key Competencies and Attributes
- Diplomatic Operator: Highly collaborative with the diplomatic communication skills needed to align divergent engineering and corporate business interests.
- Systems Thinker: Able to identify systemic process bottlenecks and solve them through automation rather than manual intervention.
- Detail-Oriented: