1

Vllm Jobs (NOW HIRING)

Experience working on high-performance server systems--you'd be just as comfortable with the internals of VLLM as you would with a complex PyTorch codebase. • Significant performance engineering ...

Senior AI Engineer

Atlanta, GA

$100K - $138K/yr

Engineer and tune LLM inference serving stacks -- primary depth in vLLM with breadth across the inference ecosystem -- for client latency, throughput, and cost targets. * Tune inference performance ...

Experience working on high-performance server systems--you'd be just as comfortable with the internals of VLLM as you would with a complex PyTorch codebase. • Significant performance engineering ...

next page

Showing results 1-20

Vllm information

How does a VLLM (Very Large Language Model) Engineer typically collaborate with data scientists and product teams during model deployment?

VLLM Engineers work closely with data scientists to understand the specific requirements and fine-tuning needs of large-scale language models. They are often responsible for integrating these models into production systems, ensuring scalability and efficiency. Collaboration with product teams is crucial to align model capabilities with user needs and to troubleshoot real-world application challenges. Frequent communication and agile workflows are common, as updates or optimizations may be needed rapidly based on feedback from both teams.

What is a VLLM and what do they do?

VLLM stands for 'Virtual Large Language Model.' In the context of AI development, VLLM professionals work with optimized inference engines for large language models, enabling faster and more efficient deployment of AI models in production environments. Their responsibilities often include integrating LLMs into applications, optimizing model performance, and ensuring scalability for real-time use cases. They may also collaborate with data scientists and engineers to manage resources and streamline AI workflows.

What is the difference between Vllm vs Data Analyst?

AspectVllmData Analyst
Required CredentialsTypically requires knowledge of machine learning, AI, and programming languages like Python or RRequires skills in statistics, Excel, SQL, and data visualization tools
Work EnvironmentOften in tech companies, research labs, or AI-focused teamsCommonly in business, finance, healthcare, and marketing sectors
Industry UsageEmerging role in AI and machine learning projectsEstablished role in data-driven decision making
Common Search/ComparisonVllm vs Data Analyst

The main difference between Vllm and Data Analyst lies in their focus and skill set. Vllm professionals specialize in AI and machine learning models, often working in tech environments, while Data Analysts focus on interpreting data to inform business decisions. Both roles require analytical skills, but Vllm roles demand programming and AI expertise, whereas Data Analysts emphasize statistical analysis and data visualization.

What are the key skills and qualifications needed to thrive as a Machine Learning Engineer working with vLLM, and why are they important?

To thrive as a Machine Learning Engineer specializing in vLLM (a high-throughput LLM inference library), you need a strong understanding of machine learning principles, deep learning frameworks, and experience with Python programming. Familiarity with tools like PyTorch, CUDA, distributed computing, and cloud platforms, as well as relevant certifications in ML or data engineering, is highly valuable. Strong problem-solving, collaboration, and communication skills are essential for optimizing model performance and integrating with cross-functional teams. These capabilities ensure effective deployment and scaling of large language models, driving innovation and efficiency in AI applications.
More about Vllm jobs
What cities are hiring for Vllm jobs? Cities with the most Vllm job openings:
What states have the most Vllm jobs? States with the most job openings for Vllm jobs include:
Infographic showing various Vllm job openings in the United States as of July 2026, with employment types broken down into 1% Internship, 96% Full Time, 1% Part Time, and 2% Contract. Highlights an 84% Physical, 3% Hybrid, and 13% Remote job distribution.

Member of Technical Staff, CI/CD Infrastructure

Inferact

San Francisco, CA • On-site

$200K - $400K/yr

Full-time

Medical, Dental, Vision, Retirement

Posted 11 days ago


Job description

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference efficient and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.
About the Role
vLLM is growing at a fast pace, and every bit of that growth lands on the CI system. More models, more hardware, more contributors, more ways for things to break. Your job is to advance the CI system so it scales with vLLM's momentum and unlocks faster development for everyone.
You'll get to:
  • Maintain and scale the compute infrastructure that powers CI, release, performance benchmark, accuracy evaluation for vLLM project, across a wide range of models and accelerators including H100/H200, (G)B200/300, AMD MI325/355X, TPU, Intel Gaudi, etc..
  • Get creative about cutting CI time-to-signal from hours to minutes
  • Make sure every corner of vLLM code base is well-tested
  • Keep vLLM releases rock-solid
  • Build out tooling that helps 3,000+ vLLM contributors move fast

Skills and Qualifications
Minimum qualifications:
  • Strong experience with Docker, Kubernetes, and containerized build or test environments.
  • Built CI/CD pipelines from scratch using GitHub Actions, Buildkite, or similar systems.
  • Familiar with CI design patterns and CI techniques: compute orchestration, handling flaky tests, dependency/environment management, caching, remote execution, test target determination, etc, test coverage, and so on.
  • Fluent in Python, Bash, Go, or similar for automation and tooling.
  • Solid fundamentals of Linux, security, networking, storage, package management,.

Bonus points for:
  • Setting up infrastructure for ML, inference, CUDA, ROCm, or accelerator-heavy workloads.
  • Running Buildkite at scale, including agents, queues, dynamic pipelines, test sharding, caching, and artifact management.
  • Operating Kubernetes clusters for CI, batch jobs, test execution, or internal developer infrastructure.
  • Managing CI/CD in large open-source project
  • Building dashboards, alerts, runbooks, or tooling for CI observability.
Logistics
  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.
  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.
  • Visa sponsorship: We sponsor visas on a case-by-case basis.
  • Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.