Job Summary:
Qumulo is a cloud data platform that manages exabytes of data for over 1,100 customers. They are seeking a Site Reliability Engineer to design and automate testing for new features, ensuring that the platform can handle demanding workloads effectively.
Responsibilities:
• Design and operationalize testing for new features: work out how customers will actually use them, how to scale-test them, and how to break them
• Automate the manual, repetitive testing our principal engineers run by hand today, using Python and our in-house frameworks on Jenkins and Argo
• Build a data-driven plan for which tests run, how often, and why, plus the framework to schedule and rerun them
• Troubleshoot build and test failures across VM instances and Qumulo-qualified hardware, from compile-time errors to integration failures
• Read cluster output and C error logs to tell a test problem from an infrastructure problem from a real bug
• Set up monitoring and alerting so problems surface early (we use OpenMetrics, Grafana, InfluxDB, and Prometheus alongside home-grown tooling)
• Help set the quality bar for releases, including a real say in what ships
• Take part in an on-call rotation for the systems your team owns
Qualifications:
Required:
• 3+ years building and operating automated testing, validation, and/or certification for complex software systems
• Strong programming ability in C
• A real breaker's instinct. You go looking for edge cases and ask 'what happens if I do this?' before anyone asks you to
• A track record of building tests yourself, not just running test plans handed to you
• Hands-on experience across both on-premises infrastructure and cloud (AWS, GCP, or Azure), with a real grasp of where each one's limits are
• Working fluency in Linux (we run Ubuntu) and Python
• A data-driven approach to deciding what to test and how often
• Troubleshoot build and test failures across VM instances and Qumulo-qualified hardware, from compile-time errors to integration failures
• Read cluster output and C error logs to tell a test problem from an infrastructure problem from a real bug
• Set up monitoring and alerting so problems surface early (we use OpenMetrics, Grafana, InfluxDB, and Prometheus alongside home-grown tooling)
• Help set the quality bar for releases, including a real say in what ships
• Take part in an on-call rotation for the systems your team owns
Preferred:
• Experience with Qumulo’s distributed file system, or parallel filesystems
• Experience with orchestration tools (Ansible, Terraform), containers, and Kubernetes
• Solid understanding of networks (routing, firewalls, security inspection devices, switch configuration)
• Storage (IOPS, Latency, read/write patterns) or protocol experience (NFS, SMB, S3)
Company:
Qumulo provides a file data platform for multi-cloud environments with large scale file data. Founded in 2012, the company is headquartered in Seattle, USA, with a team of 201-500 employees. The company is currently Growth Stage.