1

Vllm Jobs (NOW HIRING)

OR

$122K - $161K/yr

Contributing to open source communities like FlashInfer, vLLM, and SGLang What we need to see: * Masters degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)

The ideal candidate will have hands-on experience with Large Language Models (LLMs) , Vision Language Models (Vision LLMs/VLMs) , vLLM inference framework , prompt engineering , and modern Generative ...

We work directly within TensorRT-LLM, SGLang, and vLLM, building the tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public ...

next page

Showing results 1-20

Vllm information

How does a VLLM (Very Large Language Model) Engineer typically collaborate with data scientists and product teams during model deployment?

VLLM Engineers work closely with data scientists to understand the specific requirements and fine-tuning needs of large-scale language models. They are often responsible for integrating these models into production systems, ensuring scalability and efficiency. Collaboration with product teams is crucial to align model capabilities with user needs and to troubleshoot real-world application challenges. Frequent communication and agile workflows are common, as updates or optimizations may be needed rapidly based on feedback from both teams.

What is a VLLM and what do they do?

VLLM stands for 'Virtual Large Language Model.' In the context of AI development, VLLM professionals work with optimized inference engines for large language models, enabling faster and more efficient deployment of AI models in production environments. Their responsibilities often include integrating LLMs into applications, optimizing model performance, and ensuring scalability for real-time use cases. They may also collaborate with data scientists and engineers to manage resources and streamline AI workflows.

What is the difference between Vllm vs Data Analyst?

AspectVllmData Analyst
Required CredentialsTypically requires knowledge of machine learning, AI, and programming languages like Python or RRequires skills in statistics, Excel, SQL, and data visualization tools
Work EnvironmentOften in tech companies, research labs, or AI-focused teamsCommonly in business, finance, healthcare, and marketing sectors
Industry UsageEmerging role in AI and machine learning projectsEstablished role in data-driven decision making
Common Search/ComparisonVllm vs Data Analyst

The main difference between Vllm and Data Analyst lies in their focus and skill set. Vllm professionals specialize in AI and machine learning models, often working in tech environments, while Data Analysts focus on interpreting data to inform business decisions. Both roles require analytical skills, but Vllm roles demand programming and AI expertise, whereas Data Analysts emphasize statistical analysis and data visualization.

What are the key skills and qualifications needed to thrive as a Machine Learning Engineer working with vLLM, and why are they important?

To thrive as a Machine Learning Engineer specializing in vLLM (a high-throughput LLM inference library), you need a strong understanding of machine learning principles, deep learning frameworks, and experience with Python programming. Familiarity with tools like PyTorch, CUDA, distributed computing, and cloud platforms, as well as relevant certifications in ML or data engineering, is highly valuable. Strong problem-solving, collaboration, and communication skills are essential for optimizing model performance and integrating with cross-functional teams. These capabilities ensure effective deployment and scaling of large language models, driving innovation and efficiency in AI applications.
More about Vllm jobs
What cities are hiring for Vllm jobs? Cities with the most Vllm job openings:
What states have the most Vllm jobs? States with the most job openings for Vllm jobs include:
Infographic showing various Vllm job openings in the United States as of July 2026, with employment types broken down into 1% Internship, 97% Full Time, 1% Part Time, and 1% Contract. Highlights an 83% Physical, 4% Hybrid, and 13% Remote job distribution.

Member of Technical Staff -- Model Optimization and Inference (New Grad)

Nuance Labs

Seattle, WA • On-site

Full-time

Posted 11 days ago


Job description

Job Summary:
Nuance Labs is a Series A company focused on developing photorealistic AI avatars with emotional intelligence. They are seeking a Member of Technical Staff to optimize model inference for real-time applications, working on end-to-end optimization across various models and contributing to the development of internal tools to enhance performance.
Responsibilities:
• Contribute to end-to-end inference optimization across our model stack — LLMs, audio models, and diffusion-based components
• Implement and tune KV cache strategies for long-context conversations, including eviction policies, compression, and memory-efficient attention
• Work with inference serving frameworks (vLLM, SGLang, TensorRT-LLM, etc.) and extend them for our specific workloads
• Profile and benchmark end-to-end latency and throughput; identify and systematically eliminate bottlenecks
• Build internal tooling that makes optimization work faster and more rigorous — profiling viewers, end-to-end inference test harnesses, and other infrastructure that helps the team move quickly
• Accelerate diffusion model inference — consistency models, step distillation, caching strategies, and custom kernel optimizations
• Apply quantization techniques (INT8, INT4, GPTQ, AWQ, and beyond) to reduce memory footprint and increase throughput without meaningfully degrading quality
• Work closely with research and infrastructure to ensure new models ship with optimized serving from day one
Qualifications:
Required:
• BS, MS, or PhD in CS, ML, or a related field — completed or in the final stretch
• Strong fundamentals in LLM inference or ML systems — KV caching, memory layout, attention kernels, batching, or serving — picked up through coursework, research, internships, or open-source. You don’t need to have shipped at production scale yet; you do need to learn fast and go deep.
• Exposure to inference serving frameworks (vLLM, SGLang, TensorRT-LLM, or similar) — even at a research or hobby level
• Strong Python and PyTorch skills; familiarity with CUDA or Triton is a significant plus
• A systematic approach to profiling and optimization — you measure first, then optimize
• Curiosity about diffusion inference, speculative decoding, quantization, or other inference-time acceleration techniques
Preferred:
• Internship or research experience with LLM inference, ML systems, or model serving
• Contributions to open-source inference frameworks (vLLM, SGLang, TensorRT-LLM, etc.)
• CUDA / Triton kernel work, even at a research or hobby scale
• Publications or research projects in MLSys, model compression, or inference optimization
• Familiarity with multimodal or streaming inference architectures
• Experience with hard latency SLAs in any real-time system
Company:
Nuance Labs an AI research company is developing the first human foundation model that understands and displays emotion in real time. Founded in 2024, the company is headquartered in Seattle, USA, with a team of 11-50 employees. The company is currently Early Stage.