Deep experience fine-tuning open-source LLMs using HuggingFace Transformers, DeepSpeed, vLLM, FSDP, LoRA/QLoRA * Worked with both base and instruction-tuned models; familiar with SFT, RLHF, DPO ...
Quick apply
Deep experience fine-tuning open-source LLMs using HuggingFace Transformers, DeepSpeed, vLLM, FSDP, LoRA/QLoRA * Worked with both base and instruction-tuned models; familiar with SFT, RLHF, DPO ...
Quick apply
Deep experience fine-tuning open-source LLMs using HuggingFace Transformers, DeepSpeed, vLLM, FSDP, LoRA/QLoRA * Worked with both base and instruction-tuned models; familiar with SFT, RLHF, DPO ...
Palo Alto, CA · On-site
$122K - $168K/yr
Build and maintain distributed training and RL infrastructure using frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, or OpenRLHF. * Implement and debug ...
Palo Alto, CA · On-site
$122K - $168K/yr
Build and maintain distributed training and RL infrastructure using frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, or OpenRLHF. * Implement and debug ...
Santa Clara, CA · On-site
$180K - $240K/yr
Deep understanding of FSDP, and DeepSpeed. * AI Agent Orchestration: Experience building Agentic Workflows (LangGraph, AutoGen) for infrastructure automation or data curation. * Advanced Protocols:
Santa Clara, CA · On-site
$180K - $240K/yr
Deep understanding of FSDP, and DeepSpeed. * AI Agent Orchestration: Experience building Agentic Workflows (LangGraph, AutoGen) for infrastructure automation or data curation. * Advanced Protocols:
Santa Clara, CA · On-site
$121K - $167K/yr
Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT. Knowledge of distributed training and multi-GPU environments. Experience with Agentic AI frameworks such as ...
Santa Clara, CA · On-site
$121K - $167K/yr
Experience with Azure AI Foundry, Azure OpenAI, Hugging Face, DeepSpeed, and PEFT. Knowledge of distributed training and multi-GPU environments. Experience with Agentic AI frameworks such as ...
San Francisco, CA · On-site
$350K - $475K/yr
Familiarity with distributed frameworks such as PyTorch/XLA, DeepSpeed, Megatron-LM. * Experience implementing FP8, INT8, or block-floating point (MX) formats and understanding their numerical trade ...
San Francisco, CA · On-site
$350K - $475K/yr
Familiarity with distributed frameworks such as PyTorch/XLA, DeepSpeed, Megatron-LM. * Experience implementing FP8, INT8, or block-floating point (MX) formats and understanding their numerical trade ...
Sunnyvale, CA · On-site
$150K - $450K/yr
Role Overview Build and scale distributed pre-training frameworks · Set up DeepSpeed / FSDP / Megatron-LM across multi-node GPU clusters. · Create robust launch scripts, resilient checkpoints, and ...
Quick apply
Sunnyvale, CA · On-site
$150K - $450K/yr
Role Overview Build and scale distributed pre-training frameworks · Set up DeepSpeed / FSDP / Megatron-LM across multi-node GPU clusters. · Create robust launch scripts, resilient checkpoints, and ...
Sunnyvale, CA · On-site
$150K - $450K/yr
Role Overview Build and scale distributed pre-training frameworks • Set up DeepSpeed / FSDP / Megatron-LM across multi-node GPU clusters. • Create robust launch scripts, resilient checkpoints ...
Sunnyvale, CA · On-site
$150K - $450K/yr
Role Overview Build and scale distributed pre-training frameworks • Set up DeepSpeed / FSDP / Megatron-LM across multi-node GPU clusters. • Create robust launch scripts, resilient checkpoints ...
$206K - $333K/yr
Cross-functional & Community - Partner with NVIDIA, key ISVs, and OSS projects (vLLM, Triton, KServe, PyTorch/DeepSpeed, ONNX Runtime) to co-develop optimizations and upstream improvements. Support ...
Quick apply
$206K - $333K/yr
Cross-functional & Community - Partner with NVIDIA, key ISVs, and OSS projects (vLLM, Triton, KServe, PyTorch/DeepSpeed, ONNX Runtime) to co-develop optimizations and upstream improvements. Support ...
Santa Clara, CA · On-site
$122K - $168K/yr
... DeepSpeed, Megatron, etc.). • A 'Product-First' mindset: The ability to balance cutting-edge research with the deterministic requirements of L4 production vehicles. Company : XPENG is a leading ...
Santa Clara, CA · On-site
$122K - $168K/yr
... DeepSpeed, Megatron, etc.). • A 'Product-First' mindset: The ability to balance cutting-edge research with the deterministic requirements of L4 production vehicles. Company : XPENG is a leading ...
Comfortable with tuning libraries (HuggingFace Trainer, DeepSpeed, FSDP, QLoRA) * Familiarity with inference serving: vLLM, TGI, Ray Serve, Triton Inference Server Automation + Infra: * Proficient ...
Quick apply
Comfortable with tuning libraries (HuggingFace Trainer, DeepSpeed, FSDP, QLoRA) * Familiarity with inference serving: vLLM, TGI, Ray Serve, Triton Inference Server Automation + Infra: * Proficient ...
Preferred : • Building or optimizing multimodal training or data pipelines. • Experience with distributed training (DeepSpeed, FSDP, Megatron-LM, etc.). • Multimodal post-training experience ...
Preferred : • Building or optimizing multimodal training or data pipelines. • Experience with distributed training (DeepSpeed, FSDP, Megatron-LM, etc.). • Multimodal post-training experience ...
Santa Clara, CA · On-site
$122K - $168K/yr
... DeepSpeed, Megatron, etc.). • A 'Product-First' mindset: The ability to balance cutting-edge research with the deterministic requirements of L4 production vehicles. Company : XPENG is a leading ...
Santa Clara, CA · On-site
$122K - $168K/yr
... DeepSpeed, Megatron, etc.). • A 'Product-First' mindset: The ability to balance cutting-edge research with the deterministic requirements of L4 production vehicles. Company : XPENG is a leading ...
Deepspeed, Huggingface TGI, FSDP) * Experience in projects involving LLMs
Deepspeed, Huggingface TGI, FSDP) * Experience in projects involving LLMs
Mountain View, CA · On-site
$35 - $50/hr
Experience with multi-node distributed training (FSDP, DeepSpeed, Megatron-LM) * Proficiency in writing custom GPU kernels with Triton or CUDA * Experience building synthetic data pipelines for agent ...
Mountain View, CA · On-site
$35 - $50/hr
Experience with multi-node distributed training (FSDP, DeepSpeed, Megatron-LM) * Proficiency in writing custom GPU kernels with Triton or CUDA * Experience building synthetic data pipelines for agent ...
Sunnyvale, CA · On-site
$125K - $164K/yr
Required : • 3+ years of experience in ML infrastructure, MLOps, or large-scale data systems • Proven experience with distributed training (PyTorch DDP, DeepSpeed, Ray, or similar) and workflow ...
Sunnyvale, CA · On-site
$125K - $164K/yr
Required : • 3+ years of experience in ML infrastructure, MLOps, or large-scale data systems • Proven experience with distributed training (PyTorch DDP, DeepSpeed, Ray, or similar) and workflow ...
Sunnyvale, CA · On-site
$125K - $164K/yr
Required : • 3+ years of experience in ML infrastructure, MLOps, or large-scale data systems • Proven experience with distributed training (PyTorch DDP, DeepSpeed, Ray, or similar) and workflow ...
Sunnyvale, CA · On-site
$125K - $164K/yr
Required : • 3+ years of experience in ML infrastructure, MLOps, or large-scale data systems • Proven experience with distributed training (PyTorch DDP, DeepSpeed, Ray, or similar) and workflow ...
Familiarity with distributed training frameworks (e.g., PyTorch, JAX, DeepSpeed, or similar). * Experience working with large-scale training or inference infrastructure. * Understanding of memory ...
Familiarity with distributed training frameworks (e.g., PyTorch, JAX, DeepSpeed, or similar). * Experience working with large-scale training or inference infrastructure. * Understanding of memory ...
Palo Alto, CA · On-site
Deepspeed, Huggingface TGI) * Experience in turning applied research results into product components
Palo Alto, CA · On-site
Deepspeed, Huggingface TGI) * Experience in turning applied research results into product components
San Francisco, CA · On-site
$203K - $241K/yr
... DeepSpeed) • Experience developing AI products, tooling, or agents Company : Baseten provides the necessary infrastructure, tooling, and expertise to integrate AI into business operations. Founded ...
San Francisco, CA · On-site
$203K - $241K/yr
... DeepSpeed) • Experience developing AI products, tooling, or agents Company : Baseten provides the necessary infrastructure, tooling, and expertise to integrate AI into business operations. Founded ...
San Francisco, CA · On-site
Familiarity with Nvidia TensorRT-LLM, vLLM, DeepSpeed, Nvidia Triton Server etc. Experience writing custom CUDA kernels using CUDA or OpenAI Triton. MS in Computer Science, Artificial Intelligence ...
San Francisco, CA · On-site
Familiarity with Nvidia TensorRT-LLM, vLLM, DeepSpeed, Nvidia Triton Server etc. Experience writing custom CUDA kernels using CUDA or OpenAI Triton. MS in Computer Science, Artificial Intelligence ...
| Aspect | Deepspeed | Data Scientist |
|---|---|---|
| Required credentials | Knowledge of machine learning frameworks, programming skills in Python, experience with AI model training | Degree in Data Science, Statistics, Computer Science, or related fields; strong analytical skills |
| Work environment | AI research labs, tech companies, cloud computing environments | Business, tech companies, research institutions |
| Industry usage | AI model training, deep learning optimization | Data analysis, predictive modeling, business insights |
Deepspeed focuses on optimizing large-scale AI model training and deep learning performance, while Data Scientists analyze data to generate insights and build predictive models. Both roles require technical skills but serve different purposes within the AI and data ecosystem.
Full-time
Re-posted yesterday
Location: San Francisco Bay Area
Type: Full-Time
Compensation: Competitive salary + meaningful equity (founding tier)
Backed by 8VC, we're building a world-class team to tackle one of the industry’s most critical infrastructure problems.
About the Role
We’re designing the future of enterprise AI infrastructure — grounded in agents, retrieval-augmented generation (RAG), knowledge graphs, and multi-tenant governance.
We’re looking for an ML/AI Research Engineer to join our AI Lab and lead the design, training, evaluation, and optimization of agent-native AI models. You'll work at the intersection of LLMs, vector search, graph reasoning, and reinforcement learning — building the intelligence layer that sits on top of our enterprise data fabric.
This isn’t a prompt engineer role. It’s full-cycle ML: from data curation and fine-tuning to evaluation, interpretability, and deployment — with cost-awareness, alignment, and agent coordination all in scope.
Core Responsibilities
Fine-tune and evaluate open-source LLMs (e.g. LLaMA 3, Mistral, Falcon, Mixtral) for enterprise use cases with both structured and unstructured data
Build and optimize RAG pipelines using LangChain, LangGraph, LlamaIndex, or Dust — integrated with our vector DBs and internal knowledge graph
Train agent architectures (ReAct, AutoGPT, BabyAGI, OpenAgents) using enterprise task data
Develop embedding-based memory and retrieval chains with token-efficient chunking strategies
Create reinforcement learning pipelines to optimize agent behaviors (e.g. RLHF, DPO, PPO)
Establish scalable evaluation harnesses for LLM and agent performance, including synthetic evals, trace capture, and explainability tools
Contribute to model observability, drift detection, error classification, and alignment
Optimize inference latency and GPU resource utilization across cloud and on-prem environments
Desired Experience
Model Training:
Deep experience fine-tuning open-source LLMs using HuggingFace Transformers, DeepSpeed, vLLM, FSDP, LoRA/QLoRA
Worked with both base and instruction-tuned models; familiar with SFT, RLHF, DPO pipelines
Comfortable building and maintaining custom training datasets, filters, and eval splits
Understand tradeoffs in batch size, token window, optimizer, precision (FP16, bfloat16), and quantization
RAG + Knowledge Graphs:
Experience building enterprise-grade RAG pipelines integrated with real-time or contextual data
Familiar with LangChain, LangGraph, LlamaIndex, and open-source vector DBs (Weaviate, Qdrant, FAISS)
Experience grounding models with structured data (SQL, graph, metadata) + unstructured sources
Bonus: Worked with Neo4j, Puppygraph, RDF, OWL, or other semantic modeling systems
Agent Intelligence:
Experience training or customizing agent frameworks with multi-step reasoning and memory
Understand common agent loop patterns (e.g. Plan→Act→Reflect), memory recall, and tools
Familiar with self-correction, multi-agent communication, and agent ops logging
Optimization:
Strong background in token cost optimization, chunking strategies, reranking (e.g. Cohere, Jina), compression, and retrieval latency tuning
Experience running models under quantized (int4/int8) or multi-GPU settings with inference tuning (vLLM, TGI)
Preferred Tech Stack
LLM Training & Inference: HuggingFace Transformers, DeepSpeed, vLLM, FlashAttention, FSDP, LoRA
Agent Orchestration: LangChain, LangGraph, ReAct, OpenAgents, LlamaIndex
Vector DBs: Weaviate, Qdrant, FAISS, Pinecone, Chroma
Graph Knowledge Systems: Neo4j, Puppygraph, RDF, Gremlin, JSON-LD
Storage & Access: Iceberg, DuckDB, Postgres, Parquet, Delta Lake
Evaluation: OpenLLM Evals, Trulens, Ragas, LangSmith, Weight & Biases
Compute: Ray, Kubernetes, TGI, Sagemaker, LambdaLabs, Modal
Languages: Python (core), optionally Rust (for inference layers) or JS (for UX experimentation)
Soft Skills & Mindset
Startup DNA: resourceful, fast-moving, and capable of working in ambiguity
Deep curiosity about agent-based architectures and real-world enterprise complexity
Comfortable owning model performance end-to-end: from dataset to deployment
Strong instincts around explainability, safety, and continuous improvement
Enjoy pair-designing with product and UX to shape capabilities, not just APIs
Why This Role Matters
This role is foundational to our thesis: that agents + enterprise data + knowledge modeling can create intelligent infrastructure for real-world, multi-billion-dollar workflows. Your work won’t be buried in research reports — it will be productionized and activated by hundreds of users and hundreds of thousands of decisions. If this is your dream role - we would love to hear from you.