1

Deepspeed Jobs in Chicago, IL (NOW HIRING)

Senior Machine Learning Engineer (LLMs)

Chicago, IL · On-site

$126K - $166K/yr

Distributed training (FSDP, DeepSpeed, Megatron, etc.) * Inference optimization (quantization, speculative decoding, vLLM, Triton) * Experience shipping LLM features in production SaaS * Open-source ...

Senior Machine Learning Engineer (LLMs)

Chicago, IL · On-site

$126K - $166K/yr

Distributed training (FSDP, DeepSpeed, Megatron, etc.) * Inference optimization (quantization, speculative decoding, vLLM, Triton) * Experience shipping LLM features in production SaaS * Opensource ...

Senior Machine Learning Engineer (LLMs)

Chicago, IL · On-site

$126K - $166K/yr

Distributed training (FSDP, DeepSpeed, Megatron, etc.) * Inference optimization (quantization, speculative decoding, vLLM, Triton) * Experience shipping LLM features in production SaaS

Deepspeed information

What is DeepSpeed?

Deepspeed is an open-source deep learning optimization library developed by Microsoft, designed to enable distributed training of large-scale models efficiently. It helps researchers and engineers train models that are too large to fit in the memory of a single GPU by offering features like ZeRO optimization, mixed-precision training, and advanced parallelism techniques. Deepspeed is widely used in the machine learning community for its scalability and performance improvements, making it easier to train state-of-the-art models on vast datasets. The library integrates seamlessly with PyTorch and supports training on multiple GPUs and even across multiple machines.

What are some common challenges faced by engineers working with DeepSpeed and how can they be addressed?

Engineers working with DeepSpeed often encounter challenges related to optimizing large-scale model training, such as managing memory efficiency and tuning distributed training parameters. Troubleshooting issues like gradient accumulation, parallelism strategies, and ensuring compatibility with different hardware setups can be complex. Collaborating closely with data scientists, DevOps, and research teams is essential for addressing these challenges, as is staying updated with the latest DeepSpeed releases and documentation. Regular participation in code reviews and knowledge-sharing sessions can also help engineers overcome technical hurdles and continuously improve model performance.

What are the key skills and qualifications needed to thrive as a DeepSpeed engineer, and why are they important?

To thrive as a DeepSpeed Engineer, you need a solid background in machine learning, deep learning frameworks (such as PyTorch), and distributed systems, often supported by a degree in computer science or a related field. Proficiency with DeepSpeed, parallel computing libraries, and cloud platforms, along with familiarity with tools like CUDA and NCCL, is typically expected. Strong problem-solving abilities, collaboration, and adaptability are crucial soft skills for optimizing large-scale AI models and working with cross-functional teams. Mastering these skills ensures efficient development and deployment of high-performance, scalable AI solutions in demanding environments.

What is the difference between Deepspeed vs Data Scientist?

AspectDeepspeedData Scientist
Required credentialsKnowledge of machine learning frameworks, programming skills in Python, experience with AI model trainingDegree in Data Science, Statistics, Computer Science, or related fields; strong analytical skills
Work environmentAI research labs, tech companies, cloud computing environmentsBusiness, tech companies, research institutions
Industry usageAI model training, deep learning optimizationData analysis, predictive modeling, business insights

Deepspeed focuses on optimizing large-scale AI model training and deep learning performance, while Data Scientists analyze data to generate insights and build predictive models. Both roles require technical skills but serve different purposes within the AI and data ecosystem.

What job categories do people searching Deepspeed jobs in Chicago, IL look for?

The top searched job categories for Deepspeed jobs in Chicago, IL are:

Senior Machine Learning Engineer (LLMs)

Albi

Chicago, IL • On-site

$126K - $166K/yr

Full-time

Medical, Dental, Vision, Retirement, PTO

Re-posted 16 days ago


Job description

We're building deeply integrated LLMs into a real product used daily by restoration companies running thousands of jobs. This is not a "prompt engineer" role. You'll design, train, and ship domain-specific language models that automate real workflows and move real revenue.
You will:
  • Own end-to-end LLM systems: architecture, training, evals, and iteration
  • Fine-tune and extend existing models (LoRA, instruction tuning, RLHF)
  • Build and maintain data pipelines from product databases, documents, APIs, and logs
  • Ship reliable, monitored, production models with clear guardrails
  • Collaborate closely with product and engineering to turn messy real-world problems into working systems
  • Build and coordinate the AI engineering team
  • Use Claude Code as a core tool for development, refactors, tests, and experiments

This is for you if:
  • "How does this actually work under the hood?" is your default question
  • You're fine sitting with a hard problem for days and reading papers on weekends to figure it out
  • If there's something interesting to learn or solve, it doesn't matter if it's Saturday or 1 a.m., you're in
  • You build side projects nobody asked for and write cleaner code than anyone requires
  • You're quietly competitive, self-taught in at least one major skill, and think in systems
  • You're slightly allergic to meetings without a clear purpose or owner

Requirements
  • 5+ years of real world experience in ML / AI engineering
  • Proven experience training or substantially contributing to training LLMs (not just calling APIs)
  • Deep understanding of transformers, attention, and training dynamics
  • Strong Python plus PyTorch or JAX
  • Experience with large-scale data pipelines and experiment tracking
  • Hands-on fine-tuning (LoRA, instruction / SFT, RLHF or similar)
  • Comfortable using Claude Code as part of your daily workflow
  • Able to explain complex systems simply to non-technical stakeholders and go deep with experts
  • Track record of owning projects end-to-end and mentoring other engineers

Nice to have:
  • Distributed training (FSDP, DeepSpeed, Megatron, etc.)
  • Inference optimization (quantization, speculative decoding, vLLM, Triton)
  • Experience shipping LLM features in production SaaS
  • Open-source contributions or published work or patents in ML / NLP
  • Microsoft Foundry experience

Benefits
  • Competitive salary (based on experience and location)
  • Generous PTO
  • Medical, dental, and vision coverage
  • 401(k) plan
  • High ownership and autonomy over your work
  • Direct collaboration with a small team of smart, kind, motivated engineers
  • An environment that values deep work, clear thinking, and real impact
  • Regular team events and off-sites
  • Equipment and learning budget to help you do your best work and keep up with the frontier