Job Title : Site Reliability Engineer (SRE)
Location : Sunnyvale, CA/Austin, TX
Client: TCS / CL-Technology, Software and Services
Positions: 2
JD :
Hybrid Role ((3 days a week))
Job Title: SRE
Role Descriptions:
We're looking for an SRE to ensure the reliability, scalability, and performance of our production systems. You'll build automation, own observability, and drive incident response across cloud-native infrastructure.
Key Responsibilities
Design, deploy, and maintain infrastructure on AWS (EC2, EKS, S3, RDS, Lambda, IAM, VPC).
Manage and scale containerized workloads using Kubernetes (deployments, autoscaling, service mesh, config/secrets management).
Build and maintain CI/CD pipelines for automated, reliable deployments.
Implement Infrastructure as Code (Terraform/CloudFormation) for reproducible environments.
Set up monitoring, logging, and alerting (Prometheus, Grafana, CloudWatch, ELK) to ensure system observability.
Define and track SLIs/SLOs/SLAs; drive error budget-based decision making.
Lead incident response, root cause analysis, and post-mortems to improve system resilience.
Automate operational toil through scripting (Python/Bash/Go) and self-healing systems.
Collaborate with development teams to embed reliability, scalability, and security best practices into the SDLC.
Manage capacity planning, cost optimization, and performance tuning across cloud workloads.
Skills: Digital : Site Reliability Engineering (SRE)~AWS DevOps and Automation
Experience Required: 6-8
** All submissions must have LinkedIn id of Candidate**