1

Deepspeed Jobs (NOW HIRING)

Experience with distributed training frameworks, such as DeepSpeed, FSDL, Ray * Experience with inference frameworks like vLLM * Experience with large-scale data processing (Spark, Beam) and ...

New

Senior Principal AI Engineer

$128K - $177K/yr

DeepSpeed * Own multi-node launch configurations, failure recovery, and performance tuning * Memory & Performance Optimization * Apply advanced memory optimization techniques: * Activation ...

... LM, DeepSpeed, Ray, PyTorch Lightning--turning raw compute into a reliable model-building factory. Preferred : • A forward-looking perspective on co-designing algorithms for unconventional ...

Showing results 41-60

Deepspeed information

What is DeepSpeed?

Deepspeed is an open-source deep learning optimization library developed by Microsoft, designed to enable distributed training of large-scale models efficiently. It helps researchers and engineers train models that are too large to fit in the memory of a single GPU by offering features like ZeRO optimization, mixed-precision training, and advanced parallelism techniques. Deepspeed is widely used in the machine learning community for its scalability and performance improvements, making it easier to train state-of-the-art models on vast datasets. The library integrates seamlessly with PyTorch and supports training on multiple GPUs and even across multiple machines.

What are some common challenges faced by engineers working with DeepSpeed and how can they be addressed?

Engineers working with DeepSpeed often encounter challenges related to optimizing large-scale model training, such as managing memory efficiency and tuning distributed training parameters. Troubleshooting issues like gradient accumulation, parallelism strategies, and ensuring compatibility with different hardware setups can be complex. Collaborating closely with data scientists, DevOps, and research teams is essential for addressing these challenges, as is staying updated with the latest DeepSpeed releases and documentation. Regular participation in code reviews and knowledge-sharing sessions can also help engineers overcome technical hurdles and continuously improve model performance.

What are the key skills and qualifications needed to thrive as a DeepSpeed engineer, and why are they important?

To thrive as a DeepSpeed Engineer, you need a solid background in machine learning, deep learning frameworks (such as PyTorch), and distributed systems, often supported by a degree in computer science or a related field. Proficiency with DeepSpeed, parallel computing libraries, and cloud platforms, along with familiarity with tools like CUDA and NCCL, is typically expected. Strong problem-solving abilities, collaboration, and adaptability are crucial soft skills for optimizing large-scale AI models and working with cross-functional teams. Mastering these skills ensures efficient development and deployment of high-performance, scalable AI solutions in demanding environments.

What is the difference between Deepspeed vs Data Scientist?

AspectDeepspeedData Scientist
Required credentialsKnowledge of machine learning frameworks, programming skills in Python, experience with AI model trainingDegree in Data Science, Statistics, Computer Science, or related fields; strong analytical skills
Work environmentAI research labs, tech companies, cloud computing environmentsBusiness, tech companies, research institutions
Industry usageAI model training, deep learning optimizationData analysis, predictive modeling, business insights

Deepspeed focuses on optimizing large-scale AI model training and deep learning performance, while Data Scientists analyze data to generate insights and build predictive models. Both roles require technical skills but serve different purposes within the AI and data ecosystem.

More about Deepspeed jobs

What cities are hiring for Deepspeed jobs?

Cities with the most Deepspeed job openings:

What states have the most Deepspeed jobs?

States with the most job openings for Deepspeed jobs include:

Infographic showing various Deepspeed job openings in the United States as of August 2026, with employment types broken down into 1% Internship, 98% Full Time, and 1% Contract. Highlights an 81% Physical, 2% Hybrid, and 17% Remote job distribution.

Machine Learning Engineer (Junior)

Pangram

New York, NY • On-site

$135K - $150K/yr

Full-time

Posted 2 days ago

New


Job description

Pangram Labs is hiring for a strong junior Machine Learning Engineer. In this role, you will build software to support the machine learning development cycle from data generation, to training models, to deployment and monitoring production machine learning systems in real customer environments.
At Pangram, ML engineers are highly involved in the research effort, are involved in publishing research, and regularly contribute ideas and innovations to the team. However, formal research experience is not necessary. This is an in-person role in our office in Downtown Brooklyn, NYC.
Responsibilities:
  • Build robust data pipelines that mine the Internet at scale and generate millions of synthetic text examples for training detection models
  • Manage distributed infrastructure for multi-GPU LLM training
  • Profiling and optimizing training and inference code
  • Deploy efficient inference pipelines for serving LLMs at scale

Requirements:
  • B.S. or M.S. in Computer Science or related areas
  • Practical experience with deep learning: internships, undergrad or masters' level research projects in an academic lab, Kaggle competitions, or interesting side projects
  • Strong programming skills in Python and modern ML frameworks
  • Excellent understanding of transformers and LLM fundamentals
  • Comfort working across research and engineering boundaries

Nice to have
  • Experience with NVIDIA GPU programming and CUDA
  • Experience with distributed training frameworks, such as DeepSpeed, FSDL, Ray
  • Experience with inference frameworks like vLLM
  • Experience with large-scale data processing (Spark, Beam) and orchestration (Airflow)
  • Experience with MLOps and experiment tracking
  • Experience with DevOps tools
  • Familiarity with cloud-based infrastructure (AWS/GCP)