Description
The
Senior Site Reliability Engineer (SRE) will implement, secure, and operate the cloud infrastructure that supports CenCore Group's proprietary enterprise SaaS platform. This role is responsible for maintaining a scalable, highly available, secure, and reliable cloud environment as the platform grows and supports enterprise customers.
Key Responsibilities
- Manage, maintain, and improve AWS-based cloud infrastructure supporting enterprise SaaS operations.
- Operate and support Kubernetes environments, including Amazon EKS.
- Own platform reliability, scalability, availability, disaster recovery readiness, and operational resilience.
- Design and support cloud networking, load balancing, routing, traffic management, and related infrastructure components.
- Implement and maintain monitoring, alerting, logging, and observability solutions to support proactive issue detection and response.
- Establish and document operational standards, Service Level Objectives (SLOs), incident response processes, and reliability best practices.
- Partner with software engineering and product teams to improve application performance, platform stability, and deployment reliability.
- Apply security best practices across IAM, secrets management, encryption, vulnerability remediation, access controls, and production operations.
- Support production operations, troubleshoot critical issues, and participate in incident resolution as needed.
Requirements
Required Qualifications
- Professional experience supporting cloud infrastructure, site reliability, DevOps, platform engineering, or systems engineering functions.
- Hands-on experience with AWS cloud services and production cloud operations.
- Experience administering or operating Kubernetes environments.
- Working knowledge of infrastructure reliability, availability, scalability, incident response, and operational support practices.
- Experience implementing monitoring, logging, alerting, or observability tools.
- Ability to troubleshoot complex production issues and coordinate resolution across technical teams.
- Strong understanding of cloud security fundamentals, including identity and access management, encryption, secrets management, and vulnerability remediation.
- Ability to document technical processes, standards, and operational procedures.
Preferred Qualifications
- Experience with AWS services such as EKS, ALB, VPC, CloudFront, Route 53, RDS/Aurora, S3, and IAM.
- Experience with Terraform or other Infrastructure as Code tools.
- Experience with monitoring platforms such as Datadog, CloudWatch, Grafana, Prometheus, or similar tools.
- PostgreSQL administration, performance tuning, or database operations experience.
- Experience supporting enterprise SaaS, cloud-native applications, or customer-facing production platforms.
- Experience developing disaster recovery, operational readiness, or production support documentation.
Skills / Competencies
- Cloud infrastructure operations and automation
- Platform reliability, scalability, and performance optimization
- Kubernetes administration and containerized application support
- Monitoring, observability, and incident response
- Cloud security and operational risk awareness
- Technical troubleshooting and root cause analysis
- Cross-functional collaboration with engineering, product, and operations teams
- Clear technical documentation and process improvement
Work Environment and Physical Requirements
This role is primarily performed in a professional office or remote technology environment, depending on business needs and position requirements. Work involves regular use of a computer, collaboration tools, and cloud-based systems. The position may require participation in production support, incident response, or after-hours troubleshooting as needed. Physical requirements are generally sedentary and include prolonged periods of sitting, computer use, and communicating with internal teams.