You'll work hands-on with some of the most advanced models in the world - such as DeepSeek R1, GPT OSS, and other frontier architectures - to push the limits of throughput, latency, and efficiency.
You'll work hands-on with some of the most advanced models in the world - such as DeepSeek R1, GPT OSS, and other frontier architectures - to push the limits of throughput, latency, and efficiency.
CVP of Applied AI FDE
San Jose, CA · On-site
Enable customer success by deeply optimizing open-sourced models (Llama 3, DeepSeek, Mixtral) and proprietary models for our specific hardware topology, utilizing tools like vLLM and TensorRT-LLM ...
CVP of Applied AI FDE
San Jose, CA · On-site
Enable customer success by deeply optimizing open-sourced models (Llama 3, DeepSeek, Mixtral) and proprietary models for our specific hardware topology, utilizing tools like vLLM and TensorRT-LLM ...
Senior AI Systems Performance Engineer
San Jose, CA · On-site
$122K - $167K/yr
You'll work hands-on with some of the most advanced models in the world - such as DeepSeek R1, GPT OSS, and other frontier architectures - to push the limits of throughput, latency, and efficiency.
Senior AI Systems Performance Engineer
San Jose, CA · On-site
$122K - $167K/yr
You'll work hands-on with some of the most advanced models in the world - such as DeepSeek R1, GPT OSS, and other frontier architectures - to push the limits of throughput, latency, and efficiency.
Fullstack Engineer - Frontend Focus
San Francisco, CA · On-site
$120K - $180K/yr
About Inference.net We combine idle GPU capacity from around the world into a single cohesive plane of compute capable of serving models like DeepSeek and Llama 4. At any moment, 5,000+ GPUs and ...
Fullstack Engineer - Frontend Focus
San Francisco, CA · On-site
$120K - $180K/yr
About Inference.net We combine idle GPU capacity from around the world into a single cohesive plane of compute capable of serving models like DeepSeek and Llama 4. At any moment, 5,000+ GPUs and ...
AI Intern
$25 - $40/hr
Run coding LLMs (Qwen, DeepSeek, Gemini) on MaxLinear servers, connected to Bitbucket, Jira and Confluence via an MCP server * Analyze test logs and publish execution summaries and reports to ...
AI Intern
$25 - $40/hr
Run coding LLMs (Qwen, DeepSeek, Gemini) on MaxLinear servers, connected to Bitbucket, Jira and Confluence via an MCP server * Analyze test logs and publish execution summaries and reports to ...
Staff Engineer, Inference Optimizations
San Francisco, CA · Remote
$191K - $239K/yr
Tune expert gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, GLM-5 etc * Hardware & Ecosystem Mastery: Act as the subject matter expert on modern GPU families (NVIDIA/AMD) and ...
Staff Engineer, Inference Optimizations
San Francisco, CA · Remote
$191K - $239K/yr
Tune expert gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, GLM-5 etc * Hardware & Ecosystem Mastery: Act as the subject matter expert on modern GPU families (NVIDIA/AMD) and ...
Familiarity with AI tools such as ChatGPT, DeepSeek, and Claude is preferred. *Please Note: Your application will be considered for all open positions. You do not need to apply for multiple roles.
Familiarity with AI tools such as ChatGPT, DeepSeek, and Claude is preferred. *Please Note: Your application will be considered for all open positions. You do not need to apply for multiple roles.
AI Intern
$25 - $40/hr
Run coding LLMs (Qwen, DeepSeek, Gemini) on MaxLinear servers, connected to Bitbucket, Jira and Confluence via an MCP server * Analyze test logs and publish execution summaries and reports to ...
AI Intern
$25 - $40/hr
Run coding LLMs (Qwen, DeepSeek, Gemini) on MaxLinear servers, connected to Bitbucket, Jira and Confluence via an MCP server * Analyze test logs and publish execution summaries and reports to ...
AI Intern
Carlsbad, CA · On-site
$25 - $40/hr
Run coding LLMs (Qwen, DeepSeek, Gemini) on MaxLinear servers, connected to Bitbucket, Jira and Confluence via an MCP server * Analyze test logs and publish execution summaries and reports to ...
AI Intern
Carlsbad, CA · On-site
$25 - $40/hr
Run coding LLMs (Qwen, DeepSeek, Gemini) on MaxLinear servers, connected to Bitbucket, Jira and Confluence via an MCP server * Analyze test logs and publish execution summaries and reports to ...
Knowledge of how to leverage OpenAI, DeepSeek, Zapier, SQL, and automation tools to move fast. Minimal engineering skills required, but strong understanding of how to prototype and build within ...
Knowledge of how to leverage OpenAI, DeepSeek, Zapier, SQL, and automation tools to move fast. Minimal engineering skills required, but strong understanding of how to prototype and build within ...
At NVIDIA, our team focuses on improving community models like Nemotron, Llama, Gemma, DeepSeek, and Qwen. As a Product Manager for Open Models, you will work with groundbreaking technology to ...
At NVIDIA, our team focuses on improving community models like Nemotron, Llama, Gemma, DeepSeek, and Qwen. As a Product Manager for Open Models, you will work with groundbreaking technology to ...
Senior Deep Learning Architect, LLM Inference
$79 - $105.75/hr
If you're passionate about pushing the boundaries of GPU hardware and software performance and understand terms like disaggregated serving, data parallel attention, MoE, Qwen3.5, DeepSeek, GPT-OSS ...
New
Senior Deep Learning Architect, LLM Inference
$79 - $105.75/hr
If you're passionate about pushing the boundaries of GPU hardware and software performance and understand terms like disaggregated serving, data parallel attention, MoE, Qwen3.5, DeepSeek, GPT-OSS ...
New
Product Manager - Open Models
Santa Clara, CA · On-site
At NVIDIA, our team focuses on improving community models like Nemotron, Llama, Gemma, DeepSeek, and Qwen. As a Product Manager for Open Models, you will work with groundbreaking technology to ...
Product Manager - Open Models
Santa Clara, CA · On-site
At NVIDIA, our team focuses on improving community models like Nemotron, Llama, Gemma, DeepSeek, and Qwen. As a Product Manager for Open Models, you will work with groundbreaking technology to ...
Product Manager, Growth
San Francisco, CA · On-site
Knowledge of how to leverage OpenAI, DeepSeek, Zapier, SQL, and automation tools to move fast. Minimal engineering skills required, but strong understanding of how to prototype and build within ...
Product Manager, Growth
San Francisco, CA · On-site
Knowledge of how to leverage OpenAI, DeepSeek, Zapier, SQL, and automation tools to move fast. Minimal engineering skills required, but strong understanding of how to prototype and build within ...
This role is responsible for development, enablement and performance tuning of a wide variety of LLM model families, including massive scale large language models like the Llama family, DeepSeek and ...
This role is responsible for development, enablement and performance tuning of a wide variety of LLM model families, including massive scale large language models like the Llama family, DeepSeek and ...
Research Engineer - Post-Training & Small Language Models (SLMs), Healthcare AI
Costa Mesa, CA · On-site
Small language models & open-weight models Train and optimize open-weight models such as Llama, Qwen, Mistral, or DeepSeek; build specialized small language models (SLMs) for on-premise and cloud ...
Research Engineer - Post-Training & Small Language Models (SLMs), Healthcare AI
Costa Mesa, CA · On-site
Small language models & open-weight models Train and optimize open-weight models such as Llama, Qwen, Mistral, or DeepSeek; build specialized small language models (SLMs) for on-premise and cloud ...
Experience deploying and fine-tuning LLMs (e.g., DeepSeek, Llama, Claude, GPT-4) and GenAI models. * Deep knowledge of Linux, distributed systems, Kubernetes, NFS, Object Storage, and GPU ...
Experience deploying and fine-tuning LLMs (e.g., DeepSeek, Llama, Claude, GPT-4) and GenAI models. * Deep knowledge of Linux, distributed systems, Kubernetes, NFS, Object Storage, and GPU ...
CVP of Applied AI FDE
San Jose, CA · On-site
$258K/yr
Enable customer success by deeply optimizing open-sourced models (Llama 3, DeepSeek, Mixtral) and proprietary models for our specific hardware topology, utilizing tools like vLLM and TensorRT-LLM.
CVP of Applied AI FDE
San Jose, CA · On-site
$258K/yr
Enable customer success by deeply optimizing open-sourced models (Llama 3, DeepSeek, Mixtral) and proprietary models for our specific hardware topology, utilizing tools like vLLM and TensorRT-LLM.
This role is responsible for development, enablement and performance tuning of a wide variety of LLM model families, including massive scale large language models like the Llama family, DeepSeek and ...
This role is responsible for development, enablement and performance tuning of a wide variety of LLM model families, including massive scale large language models like the Llama family, DeepSeek and ...
Technical Lead Manager
San Francisco, CA · On-site
Integration with advanced reasoning models (GPT-4o, o3, Claude 3.5 Sonnet, DeepSeek) across custom function-calling, LLM fine-tuning, and hybrid vector-search RAG pipelines. * Standards: Model ...
Technical Lead Manager
San Francisco, CA · On-site
Integration with advanced reasoning models (GPT-4o, o3, Claude 3.5 Sonnet, DeepSeek) across custom function-calling, LLM fine-tuning, and hybrid vector-search RAG pipelines. * Standards: Model ...
Deepseek information
What are the key skills and qualifications needed to thrive as a deepseek engineer?
What are the typical collaboration patterns for deepseek engineers, and how do they work with cross-functional teams?
What is a deepseek engineer?
What is the difference between Deepseek vs Data Analyst?
| Aspect | Deepseek | Data Analyst |
|---|---|---|
| Required Credentials | Typically requires a background in computer science, data science, or related fields; certifications in data analysis or machine learning are common | Usually requires a degree in statistics, mathematics, or related fields; certifications like Microsoft Excel, Tableau, or SQL are beneficial |
| Work Environment | Primarily technical, involving data processing, algorithm development, and machine learning model training | Primarily analytical, involving data interpretation, reporting, and visualization |
| Employer & Industry Usage | Used in tech companies, AI firms, and research institutions focusing on machine learning and AI solutions | Used across various industries including finance, marketing, healthcare, and consulting for data-driven decision making |
Deepseek focuses on developing AI and machine learning models, requiring technical expertise in algorithms and programming. Data Analysts interpret and visualize data to support business decisions. While both roles work with data, Deepseek is more technical and research-oriented, whereas Data Analysts focus on insights and reporting.
What are popular job titles related to Deepseek jobs in California?
For Deepseek jobs in California, the most frequently searched job titles are:
What job categories do people searching Deepseek jobs in California look for?
The top searched job categories for Deepseek jobs in California are:
What cities in California are hiring for Deepseek jobs?
Cities in California with the most Deepseek job openings:

Full-time
Re-posted 24 days ago
Job description
We are seeking a talented and driven ML performance engineer to optimize and scale state-of-the-art foundation models on SambaNova's reconfigurable dataflow platform. You'll work hands-on with some of the most advanced models in the world - such as DeepSeek R1, GPT OSS, and other frontier architectures - to push the limits of throughput, latency, and efficiency. In this role, you'll bridge the gap between deep learning and systems performance, collaborating across compiler, runtime, and hardware layers to deliver world-record performance for large-scale AI inference.
Responsibilities- Bring up and optimize cutting-edge foundation models (e.g., DeepSeek, Llama, Qwen, and others) on the SambaNova platform through the SambaNova software stack.
- Profile and enhance model performance across compiler, runtime, and hardware layers to achieve SOTA throughput and latency.
- Collaborate with machine learning, compiler, runtime, and hardware teams to deliver co-designed, high-performance AI applications.
- Integrate the latest advances in model architecture, quantization, scheduling, and memory optimization from both academia and industry.
- Develop robust, scalable, and efficient end-to-end inference solutions aligned with customer needs.
- Identify performance bottlenecks and propose dataflow or scheduling optimizations for both single-node and distributed systems.
- Bachelor's or higher degree in computer science, electrical engineering, or a related field (e.g., applied mathematics, physics, or statistics).
- 3+ years of experience in one or more of the following areas:
- Deep learning model development and performance optimization
- Compiler, runtime, or kernel-level optimization
- Software-hardware co-design or systems performance tuning
- Proficiency in Python or C++, with strong foundations in algorithms, data structures, and numerical computing.
- Experience with at least one major ML framework - PyTorch, TensorFlow, or JAX.
- Demonstrated ability to analyze and optimize performance in real-world ML pipelines.
- Hands-on experience with LLM or multimodal model training and inference.
- Background in large-scale distributed training, continuous batching, and high-throughput inference systems.
- Familiarity with quantization, graph optimization, kernel fusion, and model partitioning.
- Experience with frameworks such as DeepSpeed, Megatron, vLLM, or TensorRT.
- Strong GPU programming skills (CUDA, Triton, or OpenCL); experience with cuDNN, cuBLAS, or similar libraries is a plus.
- Knowledge of memory hierarchy optimization, caching, and scheduling for large-scale model execution.
- Publication record or open-source contributions in ML systems or performance optimization is a plus.