... vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source • Follow the latest state-of-the-art in the open source and ...
... vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source • Follow the latest state-of-the-art in the open source and ...
On-prem Platform Engineer
Charlotte, NC · On-site
Must-Have Skills (Mandatory Keywords) LLM Inference & Optimization vLLM, TensorRT-LLM, Triton Inference Server, SGLang Inference optimization techniques: Continuous batching Speculative decoding KV ...
On-prem Platform Engineer
Charlotte, NC · On-site
Must-Have Skills (Mandatory Keywords) LLM Inference & Optimization vLLM, TensorRT-LLM, Triton Inference Server, SGLang Inference optimization techniques: Continuous batching Speculative decoding KV ...
Senior Software Engineer, Quantized Inference
Redmond, WA · On-site
$137K - $180K/yr
Responsibilities : • Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang) • Own model export pipelines (ModelOpt, Megatron-LM HuggingFace), ensuring quantized ...
Senior Software Engineer, Quantized Inference
Redmond, WA · On-site
$137K - $180K/yr
Responsibilities : • Implement quantized and sparse recipes in inference engines (vLLM, TRT-LLM, SGLang) • Own model export pipelines (ModelOpt, Megatron-LM HuggingFace), ensuring quantized ...
Senior Software Engineer - VLM Microservices for Neural Reconstruction
Santa Clara, CA · On-site
$143K - $189K/yr
Contribute VLM-related features to Open-Source projects like vLLM * Collaborate closely with Research and Product teams and influence our common roadmaps What we need to see: * Master's of Science in ...
Senior Software Engineer - VLM Microservices for Neural Reconstruction
Santa Clara, CA · On-site
$143K - $189K/yr
Contribute VLM-related features to Open-Source projects like vLLM * Collaborate closely with Research and Product teams and influence our common roadmaps What we need to see: * Master's of Science in ...
Stay hands-on across the AMG stack (Python, C++, CUDA, vLLM, NIXL/Dynamo, Kubernetes), contributing directly to production systems while providing technical leadership to the team. * Solve Hard ...
Senior Software Engineer - VLM Microservices for Neural Reconstruction
Santa Clara, CA · On-site
$143K - $189K/yr
Contribute VLM-related features to Open-Source projects like vLLM * Collaborate closely with Research and Product teams and influence our common roadmaps What we need to see: * Master's of Science in ...
Senior Software Engineer - VLM Microservices for Neural Reconstruction
Santa Clara, CA · On-site
$143K - $189K/yr
Contribute VLM-related features to Open-Source projects like vLLM * Collaborate closely with Research and Product teams and influence our common roadmaps What we need to see: * Master's of Science in ...
Contribute VLM-related features to Open-Source projects like vLLM * Collaborate closely with Research and Product teams and influence our common roadmaps What we need to see: * Master's of Science in ...
Contribute VLM-related features to Open-Source projects like vLLM * Collaborate closely with Research and Product teams and influence our common roadmaps What we need to see: * Master's of Science in ...
Member of Technical Staff - Model Optimization and Inference (New Grad)
Seattle, WA · On-site
$200K - $300K/yr
You've worked with vLLM, SGLang, or similar frameworks (through coursework, research, internships, or open-source) and have opinions about where they fall short. This posting is aimed at early-career ...
Member of Technical Staff - Model Optimization and Inference (New Grad)
Seattle, WA · On-site
$200K - $300K/yr
You've worked with vLLM, SGLang, or similar frameworks (through coursework, research, internships, or open-source) and have opinions about where they fall short. This posting is aimed at early-career ...
Stay hands-on across the AMG stack (Python, C++, CUDA, vLLM, NIXL/Dynamo, Kubernetes), contributing directly to production systems while providing technical leadership to the team. * Solve Hard ...
Stay hands-on across the AMG stack (Python, C++, CUDA, vLLM, NIXL/Dynamo, Kubernetes), contributing directly to production systems while providing technical leadership to the team. * Solve Hard ...
Design, optimize, and deploy highly scalable AI/ML inference systems, leveraging the latest LLM serving technologies such as vLLM, SGLang, and advanced KV Cache optimization to maximize throughput ...
Design, optimize, and deploy highly scalable AI/ML inference systems, leveraging the latest LLM serving technologies such as vLLM, SGLang, and advanced KV Cache optimization to maximize throughput ...
AI Engineer
Arlington, VA · On-site
Leverage tools and frameworks such as LangGraph, Semantic Kernel, vLLM, Ollama, and Ray for scalable AI solutions * Integrate with NVIDIA GPU ecosystems and vector databases to enhance AI performance ...
AI Engineer
Arlington, VA · On-site
Leverage tools and frameworks such as LangGraph, Semantic Kernel, vLLM, Ollama, and Ray for scalable AI solutions * Integrate with NVIDIA GPU ecosystems and vector databases to enhance AI performance ...
LLM Inference Deployment Engineer
$180K - $240K/yr
Utilize inference runtimes such as ONNX Runtime, vLLM for efficient execution. * Optimize batching, caching, and tensor parallelism to improve LLM scalability in real-time applications. * Develop and ...
LLM Inference Deployment Engineer
$180K - $240K/yr
Utilize inference runtimes such as ONNX Runtime, vLLM for efficient execution. * Optimize batching, caching, and tensor parallelism to improve LLM scalability in real-time applications. * Develop and ...
Come help craft the story for CUDA, core NVIDIA acceleration libraries like cuDNN, NCCL, NIXL, and AI frameworks like PyTorch, JAX, vLLM, and SGLang. What you'll be doing: * Own positioning and ...
Come help craft the story for CUDA, core NVIDIA acceleration libraries like cuDNN, NCCL, NIXL, and AI frameworks like PyTorch, JAX, vLLM, and SGLang. What you'll be doing: * Own positioning and ...
CUDA Libraries and Frameworks Product Marketing Manager
Santa Clara, CA · On-site
$180K/yr
Come help craft the story for CUDA, core NVIDIA acceleration libraries like cuDNN, NCCL, NIXL, and AI frameworks like PyTorch, JAX, vLLM, and SGLang. What you'll be doing: * Own positioning and ...
CUDA Libraries and Frameworks Product Marketing Manager
Santa Clara, CA · On-site
$180K/yr
Come help craft the story for CUDA, core NVIDIA acceleration libraries like cuDNN, NCCL, NIXL, and AI frameworks like PyTorch, JAX, vLLM, and SGLang. What you'll be doing: * Own positioning and ...
You've worked with vLLM, SGLang, or similar frameworks at scale and have strong opinions about where they fall short. This posting is aimed at experienced engineers and researchers who've operated at ...
You've worked with vLLM, SGLang, or similar frameworks at scale and have strong opinions about where they fall short. This posting is aimed at experienced engineers and researchers who've operated at ...
We are accelerating LLM inference across the stack and across all open source LLM frameworks like TensorRT LLM, vLLM and SGLang. With demand for AI exploding, particularly in the realm of large ...
We are accelerating LLM inference across the stack and across all open source LLM frameworks like TensorRT LLM, vLLM and SGLang. With demand for AI exploding, particularly in the realm of large ...
Responsibilities : • Own bring-up, correctness and performance of the OpenAI inference stack on AMD hardware. • Integrate internal model-serving infrastructure (e.g., vLLM, Triton) into a variety ...
Responsibilities : • Own bring-up, correctness and performance of the OpenAI inference stack on AMD hardware. • Integrate internal model-serving infrastructure (e.g., vLLM, Triton) into a variety ...
... such as vLLM Qualifications : Required : • PhD in Computer Science or a related field with 10+ years of experience in AI Serving Framework for large-scale computing, with focusing on the AI ...
... such as vLLM Qualifications : Required : • PhD in Computer Science or a related field with 10+ years of experience in AI Serving Framework for large-scale computing, with focusing on the AI ...
Senior Engineer II, AI Inference Optimization
Seattle, WA · Hybrid
$167K - $209K/yr
Familiarity with LLM serving stacks such as vLLM, TensorRT-LLM, or similar technologies * Experience building systems for inference optimization, rate limiting, routing, or workload orchestration ...
Senior Engineer II, AI Inference Optimization
Seattle, WA · Hybrid
$167K - $209K/yr
Familiarity with LLM serving stacks such as vLLM, TensorRT-LLM, or similar technologies * Experience building systems for inference optimization, rate limiting, routing, or workload orchestration ...
Senior Software Engineer - TensorRT Edge-LLM
Santa Clara, CA · On-site
$143K - $189K/yr
Preferred : • Demonstrated development experience or open-source contributions to LLM inference frameworks and libraries, such as SGLang, vLLM, or FlashInfer. • Proficiency with CUDA, including ...
Senior Software Engineer - TensorRT Edge-LLM
Santa Clara, CA · On-site
$143K - $189K/yr
Preferred : • Demonstrated development experience or open-source contributions to LLM inference frameworks and libraries, such as SGLang, vLLM, or FlashInfer. • Proficiency with CUDA, including ...
Vllm information
How does a VLLM (Very Large Language Model) Engineer typically collaborate with data scientists and product teams during model deployment?
What is a VLLM and what do they do?
What is the difference between Vllm vs Data Analyst?
| Aspect | Vllm | Data Analyst |
|---|---|---|
| Required Credentials | Typically requires knowledge of machine learning, AI, and programming languages like Python or R | Requires skills in statistics, Excel, SQL, and data visualization tools |
| Work Environment | Often in tech companies, research labs, or AI-focused teams | Commonly in business, finance, healthcare, and marketing sectors |
| Industry Usage | Emerging role in AI and machine learning projects | Established role in data-driven decision making |
| Common Search/Comparison | Vllm vs Data Analyst |
The main difference between Vllm and Data Analyst lies in their focus and skill set. Vllm professionals specialize in AI and machine learning models, often working in tech environments, while Data Analysts focus on interpreting data to inform business decisions. Both roles require analytical skills, but Vllm roles demand programming and AI expertise, whereas Data Analysts emphasize statistical analysis and data visualization.
What are the key skills and qualifications needed to thrive as a Machine Learning Engineer working with vLLM, and why are they important?
- Machine Learning Ops Engineer
- Internship Ai Engineer Salary In
- Machine Learning Algorithms
- Google Machine Learning Engineer
- Machine Learning Compiler Engineer
- Machine Learning Winter Internship
- Founding Machine Learning Engineer
- Principal Machine Learning Engineer
- Embedded Machine Learning Internship
- Contract Apple Machine Learning Engineer

Job description
Anyscale is on a mission to democratize distributed computing and make it accessible to software developers. The Distributed LLM Inference Engineer will optimize systems for large-scale ML inference, working closely with product teams and the open-source community to deliver high-performance solutions.
Responsibilities:
• Iterate very quickly with product teams to ship the end to end solutions for Batch and Online inference at high scale which will be used by open-source Ray users and customers of Anyscale
• Work across the stack integrating Ray Data and LLM engine providing optimizations achieving low cost solutions for large scale ML inference
• Integrate with Open source software like vLLM, work closely with the community to adopt these techniques in Anyscale solutions, and also contribute improvements to open source
• Follow the latest state-of-the-art in the open source and the research community, implementing and extending best practices
Qualifications:
Required:
• Familiarity with running ML inference at large scale with high throughput and low latency
• Familiarity with deep learning and deep learning frameworks (e.g. PyTorch)
• Solid understanding of distributed systems, ML inference challenges
Preferred:
• ML Systems knowledge
• Experience using Ray
• Work closely with community on LLM engines like vLLM, TensorRT-LLM
• Contributions to deep learning frameworks (PyTorch, TensorFlow)
• Contributions to deep learning compilers (Triton, TVM, MLIR)
• Prior experience working on GPUs / CUDA
Company:
Anyscale accelerates the development and productionization of any AI app, on any cloud, at any scale. Founded in 2019, the company is headquartered in San Francisco, USA, with a team of 201-500 employees. The company is currently Growth Stage.
About Anyscale
Sourced by ZipRecruiter
Industry
Software development
Company size
51 - 200 Employees
Headquarters location
San Francisco, CA, US
Year founded
2019