What's the Opportunity?Â
Reporting to the Senior Manager, Cloud Operations, the Lead Cloud Operations Engineer is a senior-level technical role responsible for leading a team of cloud engineers and hands-on technical execution in Azure/Kubernetes environments. You will be building scalable, reliable, and efficient infrastructure while mentoring junior engineers, maintaining best practices, and ensuring optimal performance, security, and cost-effectiveness in a customer-facing role.
This is a technical leadership position for someone who wants to stay close to infrastructure while growing a team. You will roll up your sleeves daily, pair with engineers on complex problems, and drive automation initiatives that reduce toil and improve platform reliability.
What You Will Do
- Lead technical strategy and best practices for Azure cloud operations and AKS deployments
- Design, implement, and maintain infrastructure-as-code using Terraform and Azure DevOps
- Manage Kubernetes deployments with ArgoCD, Helm, and GitOps patterns for auditability and compliance
- Architect and manage Azure Application Gateway, load balancing, WAF, and ingress strategies
- Oversee secrets and certificate management using Azure Key Vault and implement security best practices
- Build and maintain Azure Container Registry (ACR) pipelines and image scanning workflows
- Mentor and guide junior engineers on technical tasks, code reviews, and cloud architecture decisions
- Build automation solutions to reduce manual toil and improve operational efficiency
- Manage and troubleshoot AKS clusters, networking, storage, and identity/access control
- Own incident response, RCA documentation, and presentation to stakeholders
- Monitor infrastructure health, performance, and security posture; identify and resolve bottlenecks
- Work with peers to establish and enforce infrastructure standards and compliance controls across client environments
- Communicate proactively with customers and internal teams on infrastructure health and roadmap
- Maintain security, backup, disaster recovery, and high-availability strategies
- Drive cost optimization and FinOps practices across cloud infrastructure
- Support offshore hours and multiple time zones as needed
What You Need to Succeed
Must Haves
- 8-12 years of cloud/infrastructure engineering experience
- 5+ years of hands-on Azure cloud platform experience (VMs, networking, identity, monitoring)
- 5+ years with Kubernetes, specifically Azure Kubernetes Service (AKS) production experience
- Expert-level Terraform and Infrastructure-as-Code (IaC) practices
- Strong Azure DevOps or equivalent CI/CD platform experience
- Hands-on experience with ArgoCD or GitOps deployment patterns for Kubernetes
- Proficiency with Helm for Kubernetes package management and templating
- Azure Application Gateway configuration and management (load balancing, WAF, SSL termination)
- Azure Key Vault and secrets management implementation and best practices
- Azure Container Registry (ACR) operations, image scanning, and registry maintenance
- Proficiency in Linux administration and troubleshooting
- Automation scripting (Python, Bash/Shell) to reduce manual operational tasks
- Strong networking knowledge: TCP/IP, DNS, load balancing, VPCs, firewall rules, network troubleshooting
- Demonstrated ability to mentor and guide junior engineers
- Degree in Computer Science, Engineering, or equivalent hands-on experience
Nice to Have
- Azure certifications (AZ-104, AZ-305, AZ-700 Networking, or equivalent)
- Kubernetes certifications (CKA, CKAD)
- Cloud security scanning tools experience (Wiz, Microsoft Defender)
- Azure API Management for API governance and rate limiting
- Backup and disaster recovery solutions (Azure Backup, Site Recovery)
- FinOps and cost optimization expertise across cloud resources
- Compliance and governance frameworks (SOC2, PCI-DSS, HIPAA, regulatory reporting)
- Domain knowledge of banking, financial services, enterprise data platforms, or regulated industries
- Experience with GitOps and advanced Kubernetes patterns
- Observability and monitoring (Prometheus, Grafana, Azure Log Analytics, Application Insights)
Additional Job DetailsÂ
- Expected Salary Range: $110,000 - $135,000
- Vacancy Status: Open Position to be filled
- Mode of Work: Hybrid
- Use of AI: Zafin may use Artificial Intelligence (AI) and/or other forms of automated technology to screen and/or assess applicants for this position. Zafin will not utilize AI for conducting interviews and/or making hiring decisions.Â