1

Deepspeed Jobs in Tennessee (NOW HIRING)

Deepspeed information

What are some common challenges faced by engineers working with DeepSpeed and how can they be addressed?

Engineers working with DeepSpeed often encounter challenges related to optimizing large-scale model training, such as managing memory efficiency and tuning distributed training parameters. Troubleshooting issues like gradient accumulation, parallelism strategies, and ensuring compatibility with different hardware setups can be complex. Collaborating closely with data scientists, DevOps, and research teams is essential for addressing these challenges, as is staying updated with the latest DeepSpeed releases and documentation. Regular participation in code reviews and knowledge-sharing sessions can also help engineers overcome technical hurdles and continuously improve model performance.

What is Deepspeed?

Deepspeed is an open-source deep learning optimization library developed by Microsoft, designed to enable distributed training of large-scale models efficiently. It helps researchers and engineers train models that are too large to fit in the memory of a single GPU by offering features like ZeRO optimization, mixed-precision training, and advanced parallelism techniques. Deepspeed is widely used in the machine learning community for its scalability and performance improvements, making it easier to train state-of-the-art models on vast datasets. The library integrates seamlessly with PyTorch and supports training on multiple GPUs and even across multiple machines.

What is the difference between Deepspeed vs Data Scientist?

AspectDeepspeedData Scientist
Required credentialsKnowledge of machine learning frameworks, programming skills in Python, experience with AI model trainingDegree in Data Science, Statistics, Computer Science, or related fields; strong analytical skills
Work environmentAI research labs, tech companies, cloud computing environmentsBusiness, tech companies, research institutions
Industry usageAI model training, deep learning optimizationData analysis, predictive modeling, business insights

Deepspeed focuses on optimizing large-scale AI model training and deep learning performance, while Data Scientists analyze data to generate insights and build predictive models. Both roles require technical skills but serve different purposes within the AI and data ecosystem.

What are the key skills and qualifications needed to thrive as a DeepSpeed Engineer, and why are they important?

To thrive as a DeepSpeed Engineer, you need a solid background in machine learning, deep learning frameworks (such as PyTorch), and distributed systems, often supported by a degree in computer science or a related field. Proficiency with DeepSpeed, parallel computing libraries, and cloud platforms, along with familiarity with tools like CUDA and NCCL, is typically expected. Strong problem-solving abilities, collaboration, and adaptability are crucial soft skills for optimizing large-scale AI models and working with cross-functional teams. Mastering these skills ensures efficient development and deployment of high-performance, scalable AI solutions in demanding environments.
What are popular job titles related to Deepspeed jobs in Tennessee? For Deepspeed jobs in Tennessee, the most frequently searched job titles are:
What job categories do people searching Deepspeed jobs in Tennessee look for? The top searched job categories for Deepspeed jobs in Tennessee are:

Senior Research Scientist, HPC and AI

Oak Ridge National Laboratory

Oak Ridge, TN

$94K - $120K/yr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

Posted 15 days ago


Oak Ridge National Laboratory rating

8.8

Company rating: 8.8 out of 10

Based on 16 frontline employees who took The Breakroom Quiz

13th of 120 rated laboratories


Job description

Requisition Id 16413 

Overview:

The Analytics and AI methods at Scale (AAIMS) group in the National Center for Computational Science (NCCS) is hiring Senior Research Scientist to push the frontier of AI for science such as: scientific reasoning, federated & collaborative learning, and reinforcement learning (RL) for self-improving models on leadership-class supercomputers. You’ll help design, train, and evaluate AI systems that plan, reason, and take actions to accelerate discovery across domains (materials, chemistry, climate, fusion, biology, and more).

NCCS operates the Frontier exascale supercomputer and world-class data facilities. This role sits at the intersection of AI at scale and HPC, giving you unmatched resources to prototype new ideas, run large ablations, and translate methods into scientific impact.

Examples of Focus Areas:

  • Agentic AI for Science: Autonomous and tool-using agents for experiment design, simulation steering, data collection, and lab/compute orchestration; planning and memory; multi-agent collaboration.
  • Scientific Reasoning: Program/path-of-thought, tool-augmented and retrieval-augmented reasoning; uncertainty quantification and calibrated decisions.
  • RL & Self-Improving Models: RLHF/RLAIF, online RL, self-play, open-ended discovery, reward modeling, curriculum/active learning, data selection, iterative post-training, safety alignment and guardrails.
  • Foundation Models for Science @ Scale: Pretraining, instruction tuning, continued pretraining, Mixture-of-Experts; distributed training/inference (FSDP, DeepSpeed, Megatron-LM, tensor/sequence parallelism); scalable evaluation pipelines for reasoning and agents.
  • Federated & Collaborative Learning: Cross-silo training across institutions and facilities; privacy-preserving learning (secure aggregation, differential privacy, MPC/HE); personalization under heterogeneity; governance-aware data/model sharing; collaborative evaluation

The NCCS is the home of the world’s first exascale supercomputer Frontier. Our Leadership Computing Program (OLCF) provides world class computing facilities to applications across all computational domains and disciplines. We are an inclusive dynamic environment that welcomes those with initiative and creativity.

Major Duties and Responsibilities:

  • Develop and coordinate division activities in HPC-AI with cross-cutting initiatives in the laboratory by establishing forward-looking centers of excellence.
  • Lead and collaborate with internal and external researchers on a variety of extreme-scale AI/ML research and projects.
  • Lead in authoring peer reviewed papers, technical papers, reports, and proposals. Advance personal and staff contributions in leading professional, academic, and research organizations.
  • Advance personal and staff contributions in leading professional, academic, and research organizations.
  • Team Building & Mentorship: Provide mentorship to postdocs, students, and junior staff, fostering long-term career development.
  • Stakeholder Engagement: Effectively communicate vision, strategy, and progress to DOE sponsors, industrial partners, and international collaborators.

Basic Qualifications:

  • PhD in Computer Science, Computer Engineering, or a field closely related to the job duties of this position.
  • A minimum of 6 years of relevant research experience outside of Ph.D.
  • Demonstrated research in cross-cutting fields of HPC and/or AI.

Preferred Requirements:

  • Demonstrated leadership in conceiving, planning, and delivering large-scale HPC-AI projects with measurable scientific or technological impact.
  • Experience securing competitive funding (e.g., DOE, NSF, DARPA, industry consortia) and leading multi-institution proposals.
  • Recognition by the broader community – invited talks, keynote addresses, professional society awards, or major benchmarks.
  • Impactful open-source contributions to HPC or AI framework (e.g., Megatron-LM, DeepSpeed, Ray, Distributed RL)
  • Interdisciplinary collaboration experience – working with domain scientist (climate, material, fusion, biology) to translate methods into real discoveries.
  • Strategic vision – ability to identify long-term research directions at the intersection of AI, HPC and domain sciences.

Special Requirements:

Please submit two letters of reference when applying to this position. You may upload these directly to your application or have them sent to ORNLRecruiting@ornl.gov with the position title and number referenced in the subject line.

Instructions to upload documents to your candidate profile:

  • Login to your account via jobs.ornl.gov
  • View Profile
  • Under the My Documents section, select Add a Document

About ORNL:

As a U.S. Department of Energy (DOE) Office of Science national laboratory, ORNL has an impressive 80-year legacy of addressing the nation’s most pressing challenges. Our team is made up of over 7,000 dedicated and innovative individuals! Our goal is to create an environment where a variety of perspectives and backgrounds are valued, ensuring ORNL is known as a top choice for employment. These principles are essential for supporting our broader mission to drive scientific breakthroughs and translate them into solutions for energy, environmental, and security challenges facing the nation.

ORNL offers competitive pay and benefits programs to attract and retain individuals who demonstrate exceptional work behaviors. The laboratory provides a range of employee benefits, including medical and retirement plans and flexible work hours, to support the well-being of you and your family. Employee amenities such as on-site fitness, banking, and cafeteria facilities are also available for added convenience.

Other benefits include the following: Prescription Drug Plan, Dental Plan, Vision Plan, 401(k) Retirement Plan, Contributory Pension Plan, Life Insurance, Disability Benefits, Generous Vacation and Holidays, Parental Leave, Legal Insurance with Identity Theft Protection, Employee Assistance Plan, Flexible Spending Accounts, Health Savings Accounts, Wellness Programs, Educational Assistance, Relocation Assistance, and Employee Discounts.

If you have difficulty using the online application system or need an accommodation to apply due to a disability, please email: ORNLRecruiting@ornl.gov.

This position will remain open for a minimum of 5 days after which it will close when a qualified candidate is identified and/or hired.

We accept Word (.doc, .docx), Adobe (unsecured .pdf), Rich Text Format (.rtf), and HTML (.htm, .html) up to 5MB in size. Resumes from third party vendors will not be accepted; these resumes will be deleted and the candidates submitted will not be considered for employment.


ORNL is an equal opportunity employer. All qualified applicants, including individuals with disabilities and protected veterans, are encouraged to apply.  UT-Battelle is an E-Verify employer.


What Oak Ridge National Laboratory employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom