Job Overview
We are seeking an experienced Senior Production Engineer to support and maintain highly scalable, reliable infrastructure environments. The ideal candidate will have strong hands-on expertise in Linux systems, production operations, incident management, observability, automation, and system-level debugging.
Candidates should have deep expertise in at least one of the following areas:
- GPU-based hardware, including management, troubleshooting, and large-scale operations
- Software-Defined Networking (SDN), including OVN and OVS
- Storage technologies, including Lightbits, VAST, Pure Storage, or distributed storage platforms
Expertise across all three domains is not required. Strong, hands-on depth in at least one domain is preferred over surface-level experience across multiple areas.
Key Responsibilities
- Monitor production systems, review alerts, analyze system performance, and identify operational anomalies.
- Lead and participate in incident response, incident drills, post-mortems, and root cause analysis.
- Coordinate with cross-functional teams during critical production incidents and drive corrective actions.
- Design and implement observability solutions, including SLIs, SLOs, monitoring, and proactive alerting strategies.
- Automate routine operational processes and develop tools to improve monitoring and system reliability.
- Analyze system logs and troubleshoot complex infrastructure and production issues.
- Collaborate with software engineering teams to improve application resilience and reliability.
- Review system and infrastructure changes before deployment and recommend reliability best practices.
- Maintain high availability, scalability, and reliability across production infrastructure.
- Document operational processes, incident findings, technical solutions, and improvement plans.
Required Qualifications
- Strong experience with system architecture, design patterns, scalability, reliability, and production operations.
- Experience leading production incidents, driving root cause analysis, and coordinating cross-functional teams.
- Strong hands-on experience building observability from the ground up, including defining SLIs/SLOs and implementing monitoring and alerting strategies.
- Strong knowledge of Linux systems and Linux kernel internals, including scheduling, memory allocation, and driver subsystems.
- Proficiency in Python, Go, or another systems programming language.
- Experience with system-level debugging, including kdump and kernel panic analysis.
- Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, and Kubernetes.
- Knowledge of CI/CD tools and practices, including GitLab CI, AWX, or similar technologies.
- Strong understanding of TCP/IP, networking, and network programming.
- Experience with distributed storage systems and object, block, or file storage architectures.
- Excellent troubleshooting, analytical, and communication skills.
Preferred Qualifications
- Hands-on experience with GPU hardware management, troubleshooting, and large-scale GPU infrastructure.
- Experience with Software-Defined Networking technologies, including OVN and OVS.
- Experience with Lightbits, VAST, Pure Storage, or similar enterprise storage platforms.
- Experience supporting large-scale bare-metal or cloud infrastructure.