Senior HPC Engineer, Classified Computing
The Field Intelligence Operations Division invites candidates to apply to join the team as a Senior High Performance Computing (HPC) Engineer for Classified Computing. The role leads design, implementation, and management of HPC systems within a classified environment, collaborating with security teams, scientists, and IT leadership to ensure performance, security, and compliance.
Responsibilities
- Lead the design, deployment, and documentation of HPC systems in a classified environment.
- Guide GPU architecture strategy and identify innovation opportunities.
- Oversee cluster installation, configuration, and management, including job scheduling and resource allocation.
- Ensure compliance with security policies and conduct audits.
- Monitor performance, troubleshoot bottlenecks, and minimize downtime.
- Lead HPC projects, collaborate with scientists, and mentor junior engineers.
- Research and implement improvements to HPC infrastructure and communicate findings.
Basic Qualifications
- BS in computer science, engineering, or a related field with at least eight years of relevant experience (or equivalent).
- Eight years of HPC engineering experience focused on cluster management, parallel computing, performance optimization.
- Experience in classified environments and knowledge of NIST, DISA STIGs.
- Experience with SLURM, PBS, Moab, etc.
- Linux administration and scripting (Bash, Python, Ansible).
- Performance tuning and benchmarking tools (e.g., Ganglia, Grafana).
- Parallel programming frameworks (MPI, OpenMP, CUDA) and InfiniBand.
Preferred Qualifications
- Advanced storage solutions and parallel file systems (Lustre, GPFS, BeeGFS).
- Professional certifications (HPC Professional, Linux+, Security+).
- Leadership and project management skills.
- Strong problem solving and communication.
- Ability to manage multiple priorities in a high‑security environment.
- Federal ATO experience.
- HPC architecture and performance optimization.
- Scientific software development.
- High‑speed networking and parallel file systems.
- Troubleshooting, diagnostics, and technical support.
- Programming & Scripting: Pascal, BASIC, Delphi, Visual Basic, C, C++.
- Systems & Network Administration: Linux (RHEL/CentOS, SUSE, Debian, Ubuntu); Windows (95–10, NT‑Server 2016–2025).
- Networking: Active Directory, TCP/IP, DHCP, DNS, WINS, VPN, Citrix, Terminal Services.
- Monitoring & Management Tools: Nagios, Ganglia, HP BAC, Precise i3.
- Infrastructure & Automation: Puppet, Cobbler, Ansible, Chef, Red Hat Satellite, Kickstart, RPM optimization.
- File Systems & Archiving: Panasas (DirectFlow/panfs), DDN (GPFS), SGI DMF, StorHouse/RFS.
- HPC Tools & Job Scheduling: MOAB/MAUI, Torque, PBS Pro, Windows HPC Scheduler.
- Containerization & GPU: Docker, Kubernetes, Kubeflow, NVIDIA DGX‑1 GPU systems.
- Databases: SQL Server (2000–2008), MySQL, Zope.
- High‑speed networking: Infiniband, Mellanox, OFED, Voltaire, Force10.
- Cybersecurity in HPC/cloud environments.
- Infrastructure as Code (AWS, Terraform, Ansible, Packer).
- Supporting scientific workflows in research environments.
Special Requirements
- Must obtain and maintain a DOE Secret Compartmented Information (SCI) clearance.
- WSAP testing required; random drug testing and possible polygraph testing.
- Must obtain and retain a federal Personal Identity Verification (PIV) card; background investigation required.
Equal Opportunity Employer
ORNL is an equal‑opportunity employer. All qualified applicants, including individuals with disabilities and protected veterans, are encouraged to apply.
Benefits Summary
- Competitive pay and benefits including medical, dental, vision, 401(k), pension, life and disability insurance.
- Generous vacation, holidays, parental leave, legal insurance, employee assistance plan, flexible spending accounts, health savings accounts, wellness programs, educational assistance, relocation assistance, and employee discounts.
#J-18808-Ljbffr