1

Deepspeed Jobs in California (NOW HIRING)

Senior Machine Learning Platform Engineer

Irvine, CA · On-site

$110K - $152K/yr

Preferred : • Experience with vector databases (OpenSearch, Pinecone, Weaviate) for indexing and retrieval. • Familiarity with distributed training frameworks (Horovod, DDP/FSDP, DeepSpeed, Ray ...

Familiarity with tools like DeepSpeed, Accelerate, Unsloth, or Kubeflow for efficient large-scale model training • Training and Fine-Tuning: Expertise in fine-tuning and quantizing transformer ...

Showing results 41-60

Deepspeed information

What is DeepSpeed?

Deepspeed is an open-source deep learning optimization library developed by Microsoft, designed to enable distributed training of large-scale models efficiently. It helps researchers and engineers train models that are too large to fit in the memory of a single GPU by offering features like ZeRO optimization, mixed-precision training, and advanced parallelism techniques. Deepspeed is widely used in the machine learning community for its scalability and performance improvements, making it easier to train state-of-the-art models on vast datasets. The library integrates seamlessly with PyTorch and supports training on multiple GPUs and even across multiple machines.

What are some common challenges faced by engineers working with DeepSpeed and how can they be addressed?

Engineers working with DeepSpeed often encounter challenges related to optimizing large-scale model training, such as managing memory efficiency and tuning distributed training parameters. Troubleshooting issues like gradient accumulation, parallelism strategies, and ensuring compatibility with different hardware setups can be complex. Collaborating closely with data scientists, DevOps, and research teams is essential for addressing these challenges, as is staying updated with the latest DeepSpeed releases and documentation. Regular participation in code reviews and knowledge-sharing sessions can also help engineers overcome technical hurdles and continuously improve model performance.

What are the key skills and qualifications needed to thrive as a DeepSpeed engineer, and why are they important?

To thrive as a DeepSpeed Engineer, you need a solid background in machine learning, deep learning frameworks (such as PyTorch), and distributed systems, often supported by a degree in computer science or a related field. Proficiency with DeepSpeed, parallel computing libraries, and cloud platforms, along with familiarity with tools like CUDA and NCCL, is typically expected. Strong problem-solving abilities, collaboration, and adaptability are crucial soft skills for optimizing large-scale AI models and working with cross-functional teams. Mastering these skills ensures efficient development and deployment of high-performance, scalable AI solutions in demanding environments.

What is the difference between Deepspeed vs Data Scientist?

AspectDeepspeedData Scientist
Required credentialsKnowledge of machine learning frameworks, programming skills in Python, experience with AI model trainingDegree in Data Science, Statistics, Computer Science, or related fields; strong analytical skills
Work environmentAI research labs, tech companies, cloud computing environmentsBusiness, tech companies, research institutions
Industry usageAI model training, deep learning optimizationData analysis, predictive modeling, business insights

Deepspeed focuses on optimizing large-scale AI model training and deep learning performance, while Data Scientists analyze data to generate insights and build predictive models. Both roles require technical skills but serve different purposes within the AI and data ecosystem.

What job categories do people searching Deepspeed jobs in California look for?

The top searched job categories for Deepspeed jobs in California are:

What cities in California are hiring for Deepspeed jobs?

Cities in California with the most Deepspeed job openings:

Applied ML Engineer

Nexusflow.ai Inc.

Palo Alto, CA • On-site

Full-time

Re-posted 16 days ago


Job description

About Nexusflow.ai

Modern enterprise copilots & agents call for last-mile quality, enterprise-grade robustness and scalable operation costs, beyond simplified programming interfaces for generative AI. Nexusflow tackles this challenge, enabling enterprises to own their workflow copilots & agents stacked on top of powerful yet cost-effective, compact LLMs. We train large language models and build last-mile quality dev tooling for copilots & agents on your enterprise workflows. Our team has built the open-source LLM, NexusRaven-V2, rivaling GPT-4 in function calling with a 100X smaller model size. Our team members are also behind the scenes of Starling, the #1 ranked compact 7B chat model based on human evaluation in Chatbot Arena.

Position: Applied ML Engineers

Nexusflow is currently adding Applied ML Engineers to our team. Our Applied ML Engineers power our LLMs as well as Nexusflow's methodologies for last-mile quality tooling for copilots and agents. They build the base layer of Nexusflow's stack, contributing to tooling product and customer solutions.

Responsibilities
  • Develop LLMs targeted at powering copilots and agents built for enterprise workflows

  • Develop toolings to attain last-mile quality and robustness for copilot & agents applications (especially under low volume of manually curated data)

  • Building copilot & agent application solutions for high value customer verticals

  • Wear many hats and collaborate with the whole team for product development, deployment and customer success

Qualification Required
  • Research or industrial engineering experience in at least one of the following aspects in the context of large language model or multi-modality models: 

    • Data curation

    • Pre training

    • Instruction tuning

    • Copilots & agents building

    • Capability study and benchmarking

  • Excitement to contribute to both applied research and software engineering on productionizing the applied research outcome

Preferred
  • Working experience in fast-pace teams
  • In-depth experience in using or contributing to modern compute frameworks for LLMs (e.g. Deepspeed, Huggingface TGI) 

  • Experience in turning applied research results into product components