Applied ML Engineer
Palo Alto, CA · On-site
Deepspeed, Huggingface TGI) * Experience in turning applied research results into product components
Palo Alto, CA · On-site
Deepspeed, Huggingface TGI) * Experience in turning applied research results into product components
Palo Alto, CA · On-site
Deepspeed, Huggingface TGI) * Experience in turning applied research results into product components
Santa Clara, CA · On-site
$180 - $240/hr
Familiarity with Nvidia TensorRT-LLM, vLLM, DeepSpeed, Nvidia Triton Server etc. * Experience writing custom CUDA kernels using CUDA or OpenAI Triton. * MS in Computer Science, Artificial ...
Santa Clara, CA · On-site
$180 - $240/hr
Familiarity with Nvidia TensorRT-LLM, vLLM, DeepSpeed, Nvidia Triton Server etc. * Experience writing custom CUDA kernels using CUDA or OpenAI Triton. * MS in Computer Science, Artificial ...
Experience with multi-node distributed training (FSDP, DeepSpeed, Megatron-LM) * Proficiency in writing custom GPU kernels with Triton or CUDA * Experience building synthetic data pipelines for agent ...
Experience with multi-node distributed training (FSDP, DeepSpeed, Megatron-LM) * Proficiency in writing custom GPU kernels with Triton or CUDA * Experience building synthetic data pipelines for agent ...
Santa Clara, CA · On-site
Familiarity with Nvidia TensorRT-LLM, vLLM, DeepSpeed, Nvidia Triton Server etc. Experience writing custom CUDA kernels using CUDA or OpenAI Triton. MS in Computer Science, Artificial Intelligence ...
Santa Clara, CA · On-site
Familiarity with Nvidia TensorRT-LLM, vLLM, DeepSpeed, Nvidia Triton Server etc. Experience writing custom CUDA kernels using CUDA or OpenAI Triton. MS in Computer Science, Artificial Intelligence ...
Cupertino, CA · On-site
$128K - $177K/yr
Experience with parallel training libraries such as PyTorch Distributed (torch.distributed), DeepSpeed, or FairScale. Experience building ML models for on-device inference. Publication record at ML ...
Cupertino, CA · On-site
$128K - $177K/yr
Experience with parallel training libraries such as PyTorch Distributed (torch.distributed), DeepSpeed, or FairScale. Experience building ML models for on-device inference. Publication record at ML ...
DeepSpeed * Megatron-LM 2. Representation Learning & Method Innovation * Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems * Conduct ...
DeepSpeed * Megatron-LM 2. Representation Learning & Method Innovation * Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems * Conduct ...
Irvine, CA · On-site
$110K - $152K/yr
Preferred : • Experience with vector databases (OpenSearch, Pinecone, Weaviate) for indexing and retrieval. • Familiarity with distributed training frameworks (Horovod, DDP/FSDP, DeepSpeed, Ray ...
Irvine, CA · On-site
$110K - $152K/yr
Preferred : • Experience with vector databases (OpenSearch, Pinecone, Weaviate) for indexing and retrieval. • Familiarity with distributed training frameworks (Horovod, DDP/FSDP, DeepSpeed, Ray ...
Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT. * Knowledge of distributed training and multi-GPU environments. * Experience with Agentic AI frameworks such as ...
Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT. * Knowledge of distributed training and multi-GPU environments. * Experience with Agentic AI frameworks such as ...
Irvine, CA · On-site
$110K - $152K/yr
Preferred : • Experience with vector databases (OpenSearch, Pinecone, Weaviate) for indexing and retrieval. • Familiarity with distributed training frameworks (Horovod, DDP/FSDP, DeepSpeed, Ray ...
Irvine, CA · On-site
$110K - $152K/yr
Preferred : • Experience with vector databases (OpenSearch, Pinecone, Weaviate) for indexing and retrieval. • Familiarity with distributed training frameworks (Horovod, DDP/FSDP, DeepSpeed, Ray ...
DeepSpeed * Megatron-LM 2. Representation Learning & Method Innovation * Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems * Conduct ...
Quick apply
DeepSpeed * Megatron-LM 2. Representation Learning & Method Innovation * Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems * Conduct ...
Familiarity with Nvidia TensorRT-LLM, vLLLM, DeepSpeed, Nvidia Triton Server etc.
Familiarity with Nvidia TensorRT-LLM, vLLLM, DeepSpeed, Nvidia Triton Server etc.
Experience with frameworks such as DeepSpeed, Megatron, vLLM, or TensorRT. * Strong GPU programming skills (CUDA, Triton, or OpenCL); experience with cuDNN, cuBLAS, or similar libraries is a plus.
Experience with frameworks such as DeepSpeed, Megatron, vLLM, or TensorRT. * Strong GPU programming skills (CUDA, Triton, or OpenCL); experience with cuDNN, cuBLAS, or similar libraries is a plus.
DeepSpeed * Megatron-LM 2. Representation Learning & Method Innovation * Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems * Conduct ...
DeepSpeed * Megatron-LM 2. Representation Learning & Method Innovation * Design and improve self-supervised and multimodal learning methods for real-world autonomous driving systems * Conduct ...
Ray, DeepSpeed, HF Accelerate, FSDP) Company : Stanford University is a teaching and research university that focuses on graduate programs in law, medicine, education, and business. Founded in 1885 ...
Ray, DeepSpeed, HF Accelerate, FSDP) Company : Stanford University is a teaching and research university that focuses on graduate programs in law, medicine, education, and business. Founded in 1885 ...
Solid experience in the development and optimization of machine learning infrastructure tools like DeepSpeed, PyTorch, TensorFlow, Ray, or similar frameworks. * Strong understanding of ...
Solid experience in the development and optimization of machine learning infrastructure tools like DeepSpeed, PyTorch, TensorFlow, Ray, or similar frameworks. * Strong understanding of ...
... DeepSpeed, vLLM, FSDP, LoRA/QLoRA. • Knowledge of precision tradeoffs (FP16, bfloat16, quantization) and multi-GPU optimization. • Ability to design scalable evaluation pipelines for vision/VLMs ...
... DeepSpeed, vLLM, FSDP, LoRA/QLoRA. • Knowledge of precision tradeoffs (FP16, bfloat16, quantization) and multi-GPU optimization. • Ability to design scalable evaluation pipelines for vision/VLMs ...
Redwood City, CA · On-site
$131K - $172K/yr
Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate). You understand the nuances of mixed precision and gradient accumulation. • Infrastructure Expertise: Hands ...
Redwood City, CA · On-site
$131K - $172K/yr
Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate). You understand the nuances of mixed precision and gradient accumulation. • Infrastructure Expertise: Hands ...
Menlo Park, CA · On-site
Familiarity with tools like DeepSpeed, Accelerate, Unsloth, or Kubeflow for efficient large-scale model training • Training and Fine-Tuning: Expertise in fine-tuning and quantizing transformer ...
Menlo Park, CA · On-site
Familiarity with tools like DeepSpeed, Accelerate, Unsloth, or Kubeflow for efficient large-scale model training • Training and Fine-Tuning: Expertise in fine-tuning and quantizing transformer ...
Experience with frameworks such as DeepSpeed, Megatron, vLLM, or TensorRT. * Strong GPU programming skills (CUDA, Triton, or OpenCL); experience with cuDNN, cuBLAS, or similar libraries is a plus.
Experience with frameworks such as DeepSpeed, Megatron, vLLM, or TensorRT. * Strong GPU programming skills (CUDA, Triton, or OpenCL); experience with cuDNN, cuBLAS, or similar libraries is a plus.
Menlo Park, CA · On-site
$120 - $160/hr
Experience with distributed training on GPU clusters (SLURM, DeepSpeed, FSDP) #J-18808-Ljbffr
New
Menlo Park, CA · On-site
$120 - $160/hr
Experience with distributed training on GPU clusters (SLURM, DeepSpeed, FSDP) #J-18808-Ljbffr
New
| Aspect | Deepspeed | Data Scientist |
|---|---|---|
| Required credentials | Knowledge of machine learning frameworks, programming skills in Python, experience with AI model training | Degree in Data Science, Statistics, Computer Science, or related fields; strong analytical skills |
| Work environment | AI research labs, tech companies, cloud computing environments | Business, tech companies, research institutions |
| Industry usage | AI model training, deep learning optimization | Data analysis, predictive modeling, business insights |
Deepspeed focuses on optimizing large-scale AI model training and deep learning performance, while Data Scientists analyze data to generate insights and build predictive models. Both roles require technical skills but serve different purposes within the AI and data ecosystem.
For Deepspeed jobs in California, the most frequently searched job titles are:
The top searched job categories for Deepspeed jobs in California are:
Cities in California with the most Deepspeed job openings:
About Nexusflow.ai
Modern enterprise copilots & agents call for last-mile quality, enterprise-grade robustness and scalable operation costs, beyond simplified programming interfaces for generative AI. Nexusflow tackles this challenge, enabling enterprises to own their workflow copilots & agents stacked on top of powerful yet cost-effective, compact LLMs. We train large language models and build last-mile quality dev tooling for copilots & agents on your enterprise workflows. Our team has built the open-source LLM, NexusRaven-V2, rivaling GPT-4 in function calling with a 100X smaller model size. Our team members are also behind the scenes of Starling, the #1 ranked compact 7B chat model based on human evaluation in Chatbot Arena.
Position: Applied ML Engineers
Nexusflow is currently adding Applied ML Engineers to our team. Our Applied ML Engineers power our LLMs as well as Nexusflow's methodologies for last-mile quality tooling for copilots and agents. They build the base layer of Nexusflow's stack, contributing to tooling product and customer solutions.
ResponsibilitiesDevelop LLMs targeted at powering copilots and agents built for enterprise workflows
Develop toolings to attain last-mile quality and robustness for copilot & agents applications (especially under low volume of manually curated data)
Building copilot & agent application solutions for high value customer verticals
Wear many hats and collaborate with the whole team for product development, deployment and customer success
Research or industrial engineering experience in at least one of the following aspects in the context of large language model or multi-modality models:
Data curation
Pre training
Instruction tuning
Copilots & agents building
Capability study and benchmarking
Excitement to contribute to both applied research and software engineering on productionizing the applied research outcome
In-depth experience in using or contributing to modern compute frameworks for LLMs (e.g. Deepspeed, Huggingface TGI)
Experience in turning applied research results into product components
Sourced by ZipRecruiter
Software development
11 - 50 Employees
Daly City, CA, US
2022