Responsibilities : • Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) • Manage storage lifecycle ...
Responsibilities : • Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) • Manage storage lifecycle ...
Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) * Manage storage lifecycle operations - cluster expansion ...
Quick apply
Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) * Manage storage lifecycle operations - cluster expansion ...
Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) * Manage storage lifecycle operations - cluster expansion ...
Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data) * Manage storage lifecycle operations - cluster expansion ...
Senior Systems Engineer
Las Vegas, NV · Hybrid
$99K - $123K/yr
Experience with CNI, CRI, and CSI plugins (Portworx, Cilium, Ceph), Service Mesh (Istio), and software-defined storage (Rook-Ceph). * Automation & IaC: Advanced expertise in Ansible and Terraform to ...
Senior Systems Engineer
Las Vegas, NV · Hybrid
$99K - $123K/yr
Experience with CNI, CRI, and CSI plugins (Portworx, Cilium, Ceph), Service Mesh (Istio), and software-defined storage (Rook-Ceph). * Automation & IaC: Advanced expertise in Ansible and Terraform to ...
Senior Systems Engineer
Las Vegas, NV · On-site
$99K - $123K/yr
Experience with CNI, CRI, and CSI plugins (Portworx, Cilium, Ceph), Service Mesh (Istio), and software-defined storage (Rook-Ceph). * Automation & IaC: Advanced expertise in Ansible and Terraform to ...
Senior Systems Engineer
Las Vegas, NV · On-site
$99K - $123K/yr
Experience with CNI, CRI, and CSI plugins (Portworx, Cilium, Ceph), Service Mesh (Istio), and software-defined storage (Rook-Ceph). * Automation & IaC: Advanced expertise in Ansible and Terraform to ...
Kubernetes platforms, Bare metal provisioning systems (e.g., MAAS) • Exposure to distributed storage systems (e.g., Ceph, Weka, or similar) • Experience working in high-performance or low-latency ...
Kubernetes platforms, Bare metal provisioning systems (e.g., MAAS) • Exposure to distributed storage systems (e.g., Ceph, Weka, or similar) • Experience working in high-performance or low-latency ...
Exposure to distributed storage systems (e.g., Ceph, Weka, or similar) * Experience working in high-performance or low-latency environments What We Offer * Stock Options * 100% paid Medical, Dental ...
Exposure to distributed storage systems (e.g., Ceph, Weka, or similar) * Experience working in high-performance or low-latency environments What We Offer * Stock Options * 100% paid Medical, Dental ...
Ceph information
What are the key skills and qualifications needed to thrive in the Ceph position, and why are they important?
To thrive as a Ceph (Ceph Storage Administrator or Engineer), you need a solid background in Linux systems administration, distributed storage concepts, and networking fundamentals, typically supported by a degree in computer science or a related field. Proficiency with Ceph deployment tools, monitoring platforms, and scripting languages like Python or Bash, as well as relevant certifications such as Red Hat Certified Engineer (RHCE), are highly valuable. Strong problem-solving skills, meticulous attention to detail, and the ability to communicate technical concepts clearly with colleagues are important soft skills for this role. These competencies ensure high availability and reliability of storage systems, effective troubleshooting, and smooth collaboration within IT teams.
What are some typical daily responsibilities for a Ceph Storage Administrator?
As a Ceph Storage Administrator, your typical day involves monitoring the health and performance of the Ceph cluster, responding to alerts, and proactively addressing issues to maintain data integrity and uptime. You will also be responsible for upgrading and patching the cluster, managing user access, and performing capacity planning. Collaboration with application and infrastructure teams is key to understanding storage needs and ensuring seamless data workflows. Additionally, you may create and maintain technical documentation and participate in on-call rotations as part of a broader IT operations team. These tasks help ensure the stability and scalability of mission-critical storage systems.
What is a Ceph job?
A Ceph job typically involves managing, deploying, and maintaining Ceph, an open-source distributed storage system. Responsibilities may include configuring storage clusters, monitoring performance, troubleshooting issues, and ensuring data redundancy and scalability. Professionals in this role often work with Linux, networking, and automation tools to optimize storage solutions for enterprises and cloud environments.

Full-time
Re-posted 22 days ago
Job description
TensorWave is a company focused on delivering seamless, secure, reliable, and resilient AI compute at scale. They are seeking a Storage Operations Engineer to manage and optimize their distributed storage platforms, ensuring performance and reliability for AI and machine learning workloads.
Responsibilities:
• Operate and maintain distributed storage platforms, including Ceph (RBD, CephFS, RGW), High-performance NAS platforms (e.g., Weka, VAST Data)
• Manage storage lifecycle operations - cluster expansion, upgrades and migrations
• Monitor and maintain storage health, including capacity utilization, data distribution and balance, cluster state and recovery operations
• Analyze and troubleshoot storage performance across IOPS, throughput, and latency (including tail latency)
• Identify and remediate bottlenecks across disk subsystems, network paths (including RDMA where applicable), client access patterns
• Support incident response and root cause analysis for storage-related issues
• Ensure storage platforms meet performance expectations for GPU and Kubernetes workloads
• Operate and support Kubernetes-integrated storage - CSI drivers, StorageClasses, PersistentVolumes / PersistentVolumeClaims
• Troubleshoot storage-related issues in Kubernetes environments, including stateful workloads, performance inconsistencies, scheduling and provisioning failures
• Execute and improve automation for storage deployment and operations using Ansible, Terraform, Kubernetes manifests / Helm
• Contribute to improving monitoring and alerting, operational workflows, runbooks and documentation
• Partner with DevOps and Platform Engineering (automation and orchestration), Network Engineering (high-throughput and RDMA networking), Compute / Virtualization teams
• Help ensure end-to-end performance across compute, network, and storage layers
Qualifications:
Required:
• 4–7+ years of experience in infrastructure, systems, or storage operations
• Strong hands-on experience operating distributed storage systems in production
• Experience with Ceph (RBD, CephFS, or RGW)
• Experience with modern storage platforms such as: Weka, VAST Data, or similar high-performance systems
• Strong Linux systems knowledge
• Solid understanding of: Storage performance characteristics (IOPS, throughput, latency)
• Data replication and failure domains
• Ability to troubleshoot across: Storage systems, Network paths, Compute clients
Preferred:
• Experience supporting AI/ML or HPC workloads
• Familiarity with: NVMe-based storage architectures, RDMA or high-throughput Ethernet environments
• Experience integrating storage with Kubernetes
• Experience operating storage across multiple data centers
• Exposure to object storage and S3-compatible APIs
Company:
TensorWave is an AMD-exclusive cloud platform that leverages AMD Instinct GPUs and ROCm for high-performance AI workloads. Founded in 2023, the company is headquartered in Las Vegas, USA, with a team of 51-200 employees. The company is currently Growth Stage.