1

Ceph Storage Jobs in Tennessee (NOW HIRING)

Ceph Storage information

What are some common challenges faced by professionals working in Ceph Storage administration, and how can they be addressed?

Professionals managing Ceph Storage environments often encounter challenges such as maintaining cluster health, balancing performance and scalability, and troubleshooting hardware or network failures. These issues can be addressed by regularly monitoring cluster metrics, using Ceph's built-in tools for diagnosis, and proactively planning for capacity expansion. Collaborating closely with system administrators and network engineers is also crucial to ensure optimal integration and quick resolution of issues. Staying updated with Ceph community best practices and documentation helps administrators adapt to evolving technologies and maintain robust, high-performing storage solutions.

What is Ceph Storage?

Ceph Storage is an open-source, distributed storage system designed to provide excellent performance, reliability, and scalability. It supports object, block, and file storage in a unified platform, making it suitable for cloud infrastructure and enterprise environments. Ceph automatically manages data replication and recovery, helping to ensure high availability and fault tolerance. The system can run on commodity hardware, reducing costs and simplifying expansion as storage needs grow.

What is the difference between Ceph Storage vs Storage Engineer?

AspectCeph StorageStorage Engineer
CredentialsKnowledge of distributed storage, Linux, and storage protocolsCertifications like Cisco, CompTIA Storage+, or vendor-specific certifications
Work EnvironmentData centers, cloud environments, large-scale storage deploymentsData centers, enterprise IT teams, cloud providers
Industry UsageOpen-source storage solutions, cloud infrastructureDesign, implementation, and management of storage systems
Search/Comparison IntentTechnical understanding of storage solutionsCareer options, job roles, and skills

Ceph Storage is an open-source, distributed storage platform used to build scalable storage clusters, often in cloud and data center environments. Storage Engineers design, deploy, and maintain storage systems, including Ceph, focusing on performance and reliability. While Ceph Storage refers to the technology itself, Storage Engineers are professionals who work with Ceph and other storage solutions to meet organizational needs.

What are the key skills and qualifications needed to thrive as a Ceph Storage Engineer, and why are they important?

To thrive as a Ceph Storage Engineer, you need a strong background in Linux systems administration, networking, and distributed storage concepts, often supported by a relevant degree or certifications like Red Hat Certified Engineer (RHCE). Familiarity with tools such as Ceph, Ansible, monitoring systems like Prometheus, and scripting languages is typically required. Critical soft skills include problem-solving, attention to detail, and effective communication for troubleshooting and collaborating with IT teams. These abilities are vital to ensuring high availability, scalability, and reliability of storage solutions in enterprise environments.
What are popular job titles related to Ceph Storage jobs in Tennessee? For Ceph Storage jobs in Tennessee, the most frequently searched job titles are:
What job categories do people searching Ceph Storage jobs in Tennessee look for? The top searched job categories for Ceph Storage jobs in Tennessee are:

Senior Linux HPC Storage Engineer

ITR

Oak Ridge, TN

Full-time

Re-posted 7 days ago


Job description

  • Must be able to work a hybrid work schedule in Oak Ridge, TN
  • Must be eligible for a federal security clearance (US Citizen)

Major Duties/Responsibilities
  1. Architect, deploy, and manage large-scale HPC storage systems, including parallel file systems such as Lustre, GPFS/Spectrum Scale, BeeGFS and WEKA
  2. Design, implement, and operate large-scale Ceph storage clusters for HPC and research workloads, delivering reliable, high-performance object, block, and file storage services.
  3. Ensure the availability, performance, scalability, and security of production storage environments.
  4. Administer and optimize enterprise storage platforms such as Qumulo and NetApp in support of HPC and research workloads.
  5. Design, deploy, and maintain archival storage solutions including Spectra Logic BlackPearl and large-scale tape libraries to ensure long-term data preservation and accessibility.
  6. Integrate high-performance, enterprise, and archival storage layers into cohesive tiered storage architectures that balance cost, scalability, and performance for diverse scientific workflows.
  7. Leverage automation and monitoring solutions to minimize day-to-day maintenance while identifying opportunities to optimize system performance and management.
  8. Collaborate with researchers and technical POCs to support large data workflows and optimize I/O performance for scientific workloads.
  9. Automate storage provisioning, monitoring, and maintenance using scripting and configuration management tools.
  10. Diagnose and resolve complex storage and I/O-related issues in high-throughput, low-latency HPC environments.
  11. Evaluate emerging storage technologies (NVMe, object storage, hierarchical storage management, burst buffers) and contribute to strategic planning for future HPC systems.
  12. Work with 24/7 operations staff to streamline monitoring and troubleshooting, significantly reducing the need for off-hours support.
  13. Deliver ORNL’s mission by aligning behaviors, priorities, and interactions with our core values of Impact, Integrity, Teamwork, Safety, and Service. Promote equal opportunity by fostering a respectful workplace.
Basic Qualifications
  • A BS degree in computer science, computer engineering, information technology, information systems, science, engineering, or related discipline and 8–12 years of relevant professional experience; or an equivalent combination of education and experience.
    • Master’s degree holders: 7–10 years of relevant experience.
    • PhD holders: 4–6 years of relevant experience.
  • Five (5) or more years managing UNIX/Linux systems.
  • Demonstrated experience managing HPC storage and large-scale enterprise storage systems.
  • Three (3) or more years working with configuration management and automation tools such as Git, Jenkins, Ansible, or Puppet.
  • Proficiency with at least one scripting language (Bash, Python, Perl, etc.).
  • Strong Linux administration and advanced troubleshooting experience.
  • Experience supporting large data systems and/or HPC scientific workloads.
  • Strong desire to innovate and evaluate new technologies for HPC and storage environments.
  • Collaborative approach and ability to become a trusted advisor to research teams.
Preferred Qualifications
  • Active DOE Q, DoD Top Secret, or TS/SCI clearance is strongly preferred.
  • Solid understanding of multiple operating systems and HPC cluster technologies.
  • Experience with Rocky/CentOS/RHEL, Ubuntu, VMware.
  • Understanding of HPC job schedulers (SLURM) and user support workflows.
  • Experience with container technologies in HPC environments.
  • Experience with multiple system deployment mechanisms (Warewulf, PXEboot, Cobbler, Bright).
  • Experience with GPU clusters (NVIDIA, AMD) for AI/ML and scientific workloads.
  • Deep expertise with high-performance parallel file systems (Lustre, GPFS/Spectrum Scale, BeeGFS, WEKA).
  • Knowledge of storage networking (Infiniband, NVMe-oF, SAN/NAS architectures).
  • Familiarity with RAID, ZFS, and object storage technologies.
  • Strong background in performance monitoring, benchmarking, and I/O optimization.
  • Experience with monitoring systems such as Grafana, CheckMK, Nagios, Zabbix, Ganglia.
  • Previous experience working in a government, scientific, or other highly technical environment.
  • Strong documentation skills and ability to prepare web-based documentation.