... vLLM, SGLang, TensorRT-LLM, etc.) and extend them for our specific workloads • Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks • Build ...
... vLLM, SGLang, TensorRT-LLM, etc.) and extend them for our specific workloads • Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks • Build ...
Senior AI Software Engineer, Kernel Libraries
Santa Clara, CA · On-site
$143K - $189K/yr
Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
Senior AI Software Engineer, Kernel Libraries
Santa Clara, CA · On-site
$143K - $189K/yr
Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
$122K - $161K/yr
Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
Senior Software Engineer - AI Inference
$143K - $189K/yr
Contribute features, fixes, and optimizations upstream to vLLM/SGLang: author PRs, participate in reviews, write benchmarks/tests, and help drive designs to completion. * Implement and optimize ...
Senior Software Engineer - AI Inference
$143K - $189K/yr
Contribute features, fixes, and optimizations upstream to vLLM/SGLang: author PRs, participate in reviews, write benchmarks/tests, and help drive designs to completion. * Implement and optimize ...
Senior AI Software Engineer, Kernel Libraries
Santa Clara, CA · On-site
$143K - $189K/yr
Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
Senior AI Software Engineer, Kernel Libraries
Santa Clara, CA · On-site
$143K - $189K/yr
Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
Senior Software Engineer - AI Inference
Santa Clara, CA · On-site
$143K - $189K/yr
Contribute features, fixes, and optimizations upstream to vLLM/SGLang: author PRs, participate in reviews, write benchmarks/tests, and help drive designs to completion. * Implement and optimize ...
Senior Software Engineer - AI Inference
Santa Clara, CA · On-site
$143K - $189K/yr
Contribute features, fixes, and optimizations upstream to vLLM/SGLang: author PRs, participate in reviews, write benchmarks/tests, and help drive designs to completion. * Implement and optimize ...
Senior Software Engineer - AI Inference
$134K - $176K/yr
Contribute features, fixes, and optimizations upstream to vLLM/SGLang: author PRs, participate in reviews, write benchmarks/tests, and help drive designs to completion. * Implement and optimize ...
Senior Software Engineer - AI Inference
$134K - $176K/yr
Contribute features, fixes, and optimizations upstream to vLLM/SGLang: author PRs, participate in reviews, write benchmarks/tests, and help drive designs to completion. * Implement and optimize ...
Senior Software Engineer, Matrix Multiplication
$143K - $189K/yr
Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
Senior Software Engineer, Matrix Multiplication
$143K - $189K/yr
Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
Senior Software Engineer, Matrix Multiplication
Santa Clara, CA · On-site
$143K - $189K/yr
Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
Senior Software Engineer, Matrix Multiplication
Santa Clara, CA · On-site
$143K - $189K/yr
Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
Generative AI Engineer
$60 - $72/hr
The ideal candidate will have hands-on experience with Large Language Models (LLMs) , Vision Language Models (Vision LLMs/VLMs) , vLLM inference framework , prompt engineering , and modern Generative ...
Quick apply
Generative AI Engineer
$60 - $72/hr
The ideal candidate will have hands-on experience with Large Language Models (LLMs) , Vision Language Models (Vision LLMs/VLMs) , vLLM inference framework , prompt engineering , and modern Generative ...
Principal Engineer - Perf and Benchmarking
Bellevue, WA · On-site
$206K - $333K/yr
If MLPerf (Training & Inference), Working closely with NVIDIA (Megatron-LM, TensorRT-LLM & DGX cloud) and the open-source community (llm-d, vLLM and all popular ML frameworks) speak to you, come help ...
Principal Engineer - Perf and Benchmarking
Bellevue, WA · On-site
$206K - $333K/yr
If MLPerf (Training & Inference), Working closely with NVIDIA (Megatron-LM, TensorRT-LLM & DGX cloud) and the open-source community (llm-d, vLLM and all popular ML frameworks) speak to you, come help ...
Senior Software Engineer, AI Inference Systems
Santa Clara, CA · On-site
$143K - $189K/yr
Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features; profile and optimize the inference framework (vLLM) with methods like speculative decoding ...
Senior Software Engineer, AI Inference Systems
Santa Clara, CA · On-site
$143K - $189K/yr
Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features; profile and optimize the inference framework (vLLM) with methods like speculative decoding ...
AI Inference Performance Engineer - New College Grad 2026
Santa Clara, CA · On-site
$164K/yr
We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public ...
AI Inference Performance Engineer - New College Grad 2026
Santa Clara, CA · On-site
$164K/yr
We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public ...
AI Inference Performance Engineer
Santa Clara, CA · On-site
$164K/yr
We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public ...
AI Inference Performance Engineer
Santa Clara, CA · On-site
$164K/yr
We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public ...
AI Inference Performance Engineer - New College Grad 2026
Santa Clara, CA · On-site
$164K/yr
We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public ...
AI Inference Performance Engineer - New College Grad 2026
Santa Clara, CA · On-site
$164K/yr
We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public ...
Principal Engineer - Perf and Benchmarking
Sunnyvale, CA · On-site
$206K - $333K/yr
If MLPerf (Training & Inference), Working closely with NVIDIA (Megatron-LM, TensorRT-LLM & DGX cloud) and the open-source community (llm-d, vLLM and all popular ML frameworks) speak to you, come help ...
Quick apply
Principal Engineer - Perf and Benchmarking
Sunnyvale, CA · On-site
$206K - $333K/yr
If MLPerf (Training & Inference), Working closely with NVIDIA (Megatron-LM, TensorRT-LLM & DGX cloud) and the open-source community (llm-d, vLLM and all popular ML frameworks) speak to you, come help ...
AI Inference Performance Engineer
Santa Clara, CA · Hybrid
$164K/yr
We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public ...
AI Inference Performance Engineer
Santa Clara, CA · Hybrid
$164K/yr
We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public ...
AI Inference Engineer
San Jose, CA · On-site
Responsibilities : • Build, operate, and optimize production model-serving stacks using frameworks such as vLLM, SGLang, Triton Inference Server, TensorRT-LLM, TorchServe, or KServe • Develop and ...
AI Inference Engineer
San Jose, CA · On-site
Responsibilities : • Build, operate, and optimize production model-serving stacks using frameworks such as vLLM, SGLang, Triton Inference Server, TensorRT-LLM, TorchServe, or KServe • Develop and ...
Senior Software Engineer, Quantized Inference
Redmond, WA · On-site
$137K - $180K/yr
Responsibilities : • Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang) • Own model export pipelines (ModelOpt, Megatron-LM HuggingFace), ensuring quantized ...
Senior Software Engineer, Quantized Inference
Redmond, WA · On-site
$137K - $180K/yr
Responsibilities : • Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang) • Own model export pipelines (ModelOpt, Megatron-LM HuggingFace), ensuring quantized ...
Contribute VLM-related features to Open-Source projects like vLLM * Collaborate closely with Research and Product teams and influence our common roadmaps What we need to see: * Master's of Science in ...
Contribute VLM-related features to Open-Source projects like vLLM * Collaborate closely with Research and Product teams and influence our common roadmaps What we need to see: * Master's of Science in ...
Vllm information
How does a VLLM (Very Large Language Model) Engineer typically collaborate with data scientists and product teams during model deployment?
What is a VLLM and what do they do?
What is the difference between Vllm vs Data Analyst?
| Aspect | Vllm | Data Analyst |
|---|---|---|
| Required Credentials | Typically requires knowledge of machine learning, AI, and programming languages like Python or R | Requires skills in statistics, Excel, SQL, and data visualization tools |
| Work Environment | Often in tech companies, research labs, or AI-focused teams | Commonly in business, finance, healthcare, and marketing sectors |
| Industry Usage | Emerging role in AI and machine learning projects | Established role in data-driven decision making |
| Common Search/Comparison | Vllm vs Data Analyst |
The main difference between Vllm and Data Analyst lies in their focus and skill set. Vllm professionals specialize in AI and machine learning models, often working in tech environments, while Data Analysts focus on interpreting data to inform business decisions. Both roles require analytical skills, but Vllm roles demand programming and AI expertise, whereas Data Analysts emphasize statistical analysis and data visualization.
What are the key skills and qualifications needed to thrive as a Machine Learning Engineer working with vLLM, and why are they important?

Member of Technical Staff -- Model Optimization and Inference (New Grad)
Seattle, WA • On-site
Full-time
Posted 11 days ago
Job description
Nuance Labs is a Series A company focused on developing photorealistic AI avatars with emotional intelligence. They are seeking a Member of Technical Staff to optimize model inference for real-time applications, working on end-to-end optimization across various models and contributing to the development of internal tools to enhance performance.
Responsibilities:
• Contribute to end-to-end inference optimization across our model stack — LLMs, audio models, and diffusion-based components
• Implement and tune KV cache strategies for long-context conversations, including eviction policies, compression, and memory-efficient attention
• Work with inference serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and extend them for our specific workloads
• Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks
• Build internal tooling that makes optimization work faster and more rigorous — profiling viewers, end-to-end inference test harnesses, and other infrastructure that helps the team move quickly
• Accelerate diffusion model inference — consistency models, step distillation, caching strategies, and custom kernel optimizations
• Apply quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality
• Work closely with research and infrastructure to ensure new models ship with optimized serving from day one
Qualifications:
Required:
• BS, MS, or PhD in CS, ML, or a related field — completed or in the final stretch
• Strong fundamentals in LLM inference or ML systems — KV caching, memory layout, attention kernels, batching, or serving — picked up through coursework, research, internships, or open-source. You don’t need to have shipped at production scale yet; you do need to learn fast and go deep.
• Exposure to inference serving frameworks (vLLM, SGLang, TensorRT-LLM, or similar) — even at a research or hobby level
• Strong Python and PyTorch skills; familiarity with CUDA or Triton is a significant plus
• A systematic approach to profiling and optimization — you measure first, then optimize
• Curiosity about diffusion inference, speculative decoding, quantization, or other inference-time acceleration techniques
Preferred:
• Internship or research experience with LLM inference, ML systems, or model serving
• Contributions to open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.)
• CUDA / Triton kernel work, even at a research or hobby scale
• Publications or research projects in MLSys, model compression, or inference optimization
• Familiarity with multimodal or streaming inference architectures
• Experience with hard latency SLAs in any real-time system
Company:
Nuance Labs an AI research company is developing the first human foundation model that understands and displays emotion in real time. Founded in 2024, the company is headquartered in Seattle, USA, with a team of 11-50 employees. The company is currently Early Stage.