REQUIRED QUALIFICATIONS
7+ years of experience: experience in operating large scale applications/web services and cloud-native apps using technologies like React, Angular, Spring Boot, REST API, JPA, and other tech stacks (open-source and proprietary)
3+ years of experience with agile development
2+ years of experience with:
Build and deploying services in CI/CD with using tools (like GitHub, CircleCI, Harness, Jenkins, GitLab)
Containerization and orchestration technologies (e.g., Docker, Kubernetes, etc.)
Working in cloud platforms like GCP, Azure, AWS, etc.
EDUCATION
Bachelor’s degree or, equivalent experience (HS diploma + 4 years relevant experience)
KEY RESPONSIBLITIES
Establish performance baseline, capacity thresholds, correlate events, and define monitoring/alerting criteria
Develop automated solutions to address potential problems before they result in a service interruption
Provide impact assessment and mitigation plan for changes going into the production environment
Investigate root cause of severe and systemic outages, identify corrective actions and apply across the enterprise
Develop availability measures that align with consumer experience to accurately assess the usability of crucial services
Build capacity models to baseline transactional load compared to resource performance and leverage data to predict overall system capacity while automating load placement to avoid outages
Identify thresholds for all critical links in the data path to quickly isolate where imbalances may result in potential outages
Analyze failure points in services to model risk level and resolution steps if failure occurs.
Assist in driving architecture enhancements into system to mitigate potential failure points.
Programmatically monitor for and remediate configuration drift of critical devices
Develop response plans to potential failure points and evaluate effectiveness during planned tests
Perform comprehensive operational health checks of the entire services to identify areas of concern and track activities to drive improvements at all levels of the architecture
Provide technical coaching and direction to more junior teammates
Desired skills
Excellent knowledge of common operating systems (Unix/Linux, Windows)
Strong oral and written communication skills.
Demonstrated experience scripting or developing software and services for the cloud (Python, Go, Java, Node.js, etc)
Experience managing version control systems such as Git
Experience deploying and managing infrastructure on public clouds such as AWS or Azure.
Good knowledge of network protocols: TCP/IP, FTP, syslog etc.
Experience using an automated configuration management system (Terraform, Chef, Ansible, etc.)
Strong organizational and project management skills
Strong analytical and problem resolution skills
Experience with configuring, customizing, and extending monitoring tools (Splunk, AppD etc.)
Excellent knowledge of TCP/IP networking, and inter-networking technologies (routing/switching, proxy, firewall, load balancing etc.)
DevOps Certification a plus
Leadership
Proactively engages with cross-functional teams to resolve issues and design solutions using critical thinking and analytics skills and best practices by actively incorporating input from various sources
Strong analytical and strong problem solving skills - effectively evaluates information/data to make decisions; anticipates obstacles and develops plans to resolve
Continuous improvement oriented – actively generates process improvements; champions and drives change initiatives
Ability to deliver results in a rapidly changing dynamic environment.
Provide technical leadership and guidance to the engineering team, driving technical excellence, establishing web best practices, and ensuring the use of industry-standard methodologies and technologies
Troubleshooting and Guidance: Apply deep technical expertise to troubleshoot issues, provide guidance to the development team, and drive the resolution of technical challenges.