2

Remote Hpc System Engineer Jobs in Tennessee (NOW HIRING)

Infrastructure Platform Engineer

Oak Ridge, TN · On-site +1

$102K - $134K/yr

... remote Overview: The High-Performance Computing Systems Section within a large Research and Development Facility with the Department of Energy is seeking an HPC Infrastructure Platform Engineer to ...

New

Requirements Management and System Engineer

Oak Ridge, TN · On-site +1

$97K - $159K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

The Management and Systems Engineer's duties include: * Assists in the development and management ... remote" positions, salary to be offered is geographic dependent. Within the range, individual pay ...

Senior Systems Engineer - Remote

Brentwood, TN · On-site +1

$80K - $93K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

This is a Full Time, remote, Senior Integrated Device Engineer role. What You'll Do * Implements ... Has technical experience related to studying and analyzing system's needs, systems development ...

This position can be remote, but it'd require you to come on-site twice a year. * Must be eligible ... E/Systems Administrator/Systems Engineer * Bachelor's degree or an equivalent combination of ...

Senior Process Engineer Energy - Remote

Oak Ridge, TN · On-site +1

$90K - $173K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Overview We are seeking a Senior Process Engineer - Remote. This remote position is based out of ... This role will support the overall design and analysis of mechanical components and system designs ...

Senior Process Engineer Energy - Remote

Oak Ridge, TN · On-site +1

$90K - $173K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Overview We are seeking a Senior Process Engineer - Remote. This remote position is based out of ... This role will support the overall design and analysis of mechanical components and system designs ...

Electrical Engineer - Remote/Hybrid

Oak Ridge, TN · On-site +1

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

As an Electrical Engineer in the Oak Ridge office, you will perform a variety of electrical ... Knowledge of electrical systems is required, including power system design, load centers, fire ...

Lead Power Systems Engineer 2

Chattanooga, TN · On-site +1

$134K - $205K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... working remote from home. We are seeking a highly skilled and experienced Lead Electrical ... Conduct detailed electrical analysis utilizing time domain techniques to evaluate system ...

Senior Power Systems Engineer 2

Chattanooga, TN · On-site +1

$100K - $152K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... working remote from home. We are seeking a highly skilled and experienced Senior Electrical ... Conduct detailed electrical analysis utilizing time domain techniques to evaluate system ...

Senior Power Systems Engineer 2

Chattanooga, TN · On-site +1

$100K - $152K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... working remote from home. We are seeking a highly skilled and experienced Senior Electrical ... Conduct detailed electrical analysis utilizing time domain techniques to evaluate system ...

next page

Showing results 1-20

Remote Hpc System Engineer information

What are the key skills and qualifications needed to thrive as a remote HPC system engineer?

To thrive as a Remote HPC System Engineer, you need expertise in Linux system administration, parallel computing, networking, and a degree in computer science or related field. Familiarity with job schedulers (like Slurm), cluster management tools, scripting languages (such as Python or Bash), and certifications like CompTIA Linux+ or Red Hat Certified Engineer are highly valuable. Strong problem-solving abilities, effective communication, and self-motivation are essential soft skills for remote collaboration and troubleshooting. These skills ensure the reliable operation, optimization, and scalability of HPC systems in distributed environments.

What are some common challenges faced by remote HPC system engineers, and how can they be managed effectively?

Remote HPC System Engineers often encounter challenges such as troubleshooting complex hardware or software issues without physical access, ensuring seamless system performance, and coordinating with geographically dispersed teams. These can be managed by leveraging strong remote monitoring tools, maintaining clear documentation, and establishing effective communication channels with on-site staff. Proactively scheduling regular system health checks and participating in virtual team meetings can also help address problems quickly and maintain high system reliability.

What is the difference between Remote Hpc System Engineer vs Remote Cloud Infrastructure Engineer?

AspectRemote Hpc System EngineerRemote Cloud Infrastructure Engineer
CredentialsTypically requires Linux certifications, HPC-specific trainingOften requires cloud platform certifications (AWS, Azure, GCP)
Work EnvironmentHigh-performance computing clusters, research labsCloud platforms, data centers, virtualized environments
Industry UsageResearch, scientific computing, academiaTech, finance, enterprise IT
Search/Comparison IntentUnderstanding HPC-specific roles vs cloud rolesComparing on-premise HPC vs cloud infrastructure

The Remote Hpc System Engineer focuses on managing and optimizing high-performance computing clusters, often in research or scientific environments. In contrast, the Remote Cloud Infrastructure Engineer specializes in designing and maintaining cloud-based infrastructure across various industries. While both roles require technical expertise in system management, their environments and certifications differ, catering to distinct operational needs.

What is a remote HPC system engineer?

Remote HPC (High Performance Computing) System Engineers are IT professionals who design, implement, manage, and troubleshoot HPC systems and clusters from a remote location. They work with advanced computing infrastructure that supports scientific research, complex simulations, and large-scale data processing. Their responsibilities include configuring hardware and software, monitoring system performance, ensuring security, and providing technical support to users, all while working off-site. This role requires strong expertise in HPC technologies, operating systems like Linux, networking, and scripting, as well as effective communication skills for collaborating with distributed teams.

What are popular job titles related to Remote Hpc System Engineer jobs in Tennessee?

For Remote Hpc System Engineer jobs in Tennessee, the most frequently searched job titles are:

What job categories do people searching Remote Hpc System Engineer jobs in Tennessee look for?

The top searched job categories for Remote Hpc System Engineer jobs in Tennessee are:

What cities in Tennessee are hiring for Remote Hpc System Engineer jobs?

Cities in Tennessee with the most Remote Hpc System Engineer job openings:

Infographic showing various Remote Hpc System Engineer job openings in Tennessee as of August 2026, with employment types broken down into 1% As Needed, 82% Full Time, 15% Part Time, and 2% Contract. Highlights an 90% Physical, 4% Hybrid, and 6% Remote job distribution.

Infrastructure Platform Engineer

ITR

Oak Ridge, TN • On-site, Remote

$102K - $134K/yr

Full-time

Posted 2 days ago

New


Job description

  • Must be eligible for a federal security clearance (U.S. citizen)
  • On-site or Hybrid preferred (Oak Ridge, TN), but will consider remote

Overview:
The High-Performance Computing Systems Section within a large Research and Development Facility with the Department of Energy is seeking an HPC Infrastructure Platform Engineer to join the HPC Infrastructure group. The preferred candidate will possess commensurate knowledge, skills, and abilities in addition to relevant education, certifications, experience, and demonstrated ability to work as a member of a team.
This group provides state-of-the-art computational and data science infrastructure coupled with dedicated technical and scientific professionals tackling large-scale problems across a broad range of scientific domains for accelerating scientific discovery and engineering advances. This group hosts the Oak Ridge Leadership Computing Facility (OLCF), one of the Department of Energy’s (DOE) National User Facilities, which operates Frontier, the nation’s first exascale supercomputer.
Major Duties/Responsibilities:
Linux Administration:
  • Deploy, configure, and manage HPC-scale services in a Linux environment, primarily Red Hat and Rocky
  • Perform regular patching, updates, and backups
  • Monitor systems using tools like Nagios and Grafana
  • Respond to and assist in troubleshooting issues
Kubernetes Administration and Automation:
  • Build and maintain foundational internal platforms and tools to enable the HPC Infrastructure team to reliably deploy, monitor, and scale applications
  • Design standardized and automated workflow patterns, build and maintain CI/CD pipelines
  • Offer self-service, excellent documentation, and assistance to HPC Infrastructure group members for efficient consumption of platform services
  • Develop, maintain, and review high-quality code for internal tools using programming languages such as Python, Golang, or Rust
  • Define policies and procedures for automation and configuration management for the team and organization as a whole
Project Management and Leadership:
  • Lead small Infrastructure projects through the project lifecycle
  • Mentor and train junior staff, creating training documentation, holding knowledge sharing sessions, and fostering skill growth throughout the team
  • Propose and implement improvements to existing Infrastructure systems as well as new systems, processes, and procedures
Basic Qualifications:
  • Bachelor’s degree in computer science or closely related field and a minimum of 5 years of experience in Linux systems and Kubernetes platform administration, or a master’s degree and a minimum of 4 years of experience in Linux systems and Kubernetes platform administration
  • An equivalent combination of education and experience will be considered
Preferred Qualifications:
  • Excellent interpersonal/communication skills and the ability to work within a team
  • Strong experience designing, building, and maintaining Kubernetes platform tools
  • Strong working knowledge of Linux system fundamentals and common network protocols
  • Programming and scripting skills in common languages such as Python and bash
  • Understanding of versioning and code review tools like GitHub and GitLab
  • Experience implementing and supporting highly-available systems and services
  • Experience with configuration management tools such as Puppet or Ansible
  • Experience deploying and maintaining virtual environments using VMware
  • Experience deploying, maintaining, and troubleshooting a variety of infrastructure services such as OpenLDAP, DNS, DHCP, etc.
  • Ability to plan, prioritize, and complete assigned projects with minimal supervision