Hands-on experience with modern ML inference and serving frameworks such as vLLM, TensorRT-LLM, SGLang, Triton Inference Server, TGI, or equivalent technologies. * Experience managing and processing ...
Hands-on experience with modern ML inference and serving frameworks such as vLLM, TensorRT-LLM, SGLang, Triton Inference Server, TGI, or equivalent technologies. * Experience managing and processing ...
Hands-on experience with modern ML inference and serving frameworks such as vLLM, TensorRT-LLM, SGLang, Triton Inference Server, TGI, or equivalent technologies. * Experience managing and processing ...
Hands-on experience with modern ML inference and serving frameworks such as vLLM, TensorRT-LLM, SGLang, Triton Inference Server, TGI, or equivalent technologies. * Experience managing and processing ...
Member of Technical Staff, Inference
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, Inference
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, TPU Performance Engineering
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, TPU Performance Engineering
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, Developer Relations
San Francisco, CA · On-site
$200K - $400K/yr
Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM ...
Member of Technical Staff, Developer Relations
San Francisco, CA · On-site
$200K - $400K/yr
Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM ...
Member of Technical Staff, Cloud Orchestration
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, Cloud Orchestration
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, AMD GPU Performance Engineering
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, AMD GPU Performance Engineering
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, Kernel Engineering
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, Kernel Engineering
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Head of Engineering
San Francisco, CA · On-site
Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM ...
Head of Engineering
San Francisco, CA · On-site
Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM ...
GPU Software Engineer/GPU Architect
San Jose, CA · On-site
$164K - $202K/yr
Integrate GPU kernels with vLLM, SGLang , and other inference servers * Build highperformance components in C++ and Python * Support AI frameworks such as PyTorch and TensorFlow * Optimize multiGPU ...
GPU Software Engineer/GPU Architect
San Jose, CA · On-site
$164K - $202K/yr
Integrate GPU kernels with vLLM, SGLang , and other inference servers * Build highperformance components in C++ and Python * Support AI frameworks such as PyTorch and TensorFlow * Optimize multiGPU ...
Senior AI Software Engineer, Kernel Libraries
Santa Clara, CA · On-site
$143K - $189K/yr
PyTorch, JAX, TensorFlow, ONNX, etc) and ideally inference engines and runtimes such as vLLM, SGLang, and MLC. • Strong Python and C/C++ programming skills Preferred : • Background in domain ...
Senior AI Software Engineer, Kernel Libraries
Santa Clara, CA · On-site
$143K - $189K/yr
PyTorch, JAX, TensorFlow, ONNX, etc) and ideally inference engines and runtimes such as vLLM, SGLang, and MLC. • Strong Python and C/C++ programming skills Preferred : • Background in domain ...
AI Inference Performance Engineer - New College Grad 2026
Santa Clara, CA · On-site
$120 - $160/hr
We work within TensorRT-LLM, SGLang, and vLLM, building tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public accountability.
AI Inference Performance Engineer - New College Grad 2026
Santa Clara, CA · On-site
$120 - $160/hr
We work within TensorRT-LLM, SGLang, and vLLM, building tools that evaluate serving performance at scale. This team sits at the intersection of GPU performance engineering and public accountability.
Sr. Software Engineer - AI Triton Kernels
San Jose, CA · On-site
$180 - $280/hr
Triton is a widely adopted language and compiler for high-performance GPU kernels, powering major AI frameworks such as PyTorch, vLLM, and SGLang. As AI workloads increasingly rely on Triton-based ...
Sr. Software Engineer - AI Triton Kernels
San Jose, CA · On-site
$180 - $280/hr
Triton is a widely adopted language and compiler for high-performance GPU kernels, powering major AI frameworks such as PyTorch, vLLM, and SGLang. As AI workloads increasingly rely on Triton-based ...
Member of Technical Staff, Performance and Scale
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Member of Technical Staff, Performance and Scale
San Francisco, CA · On-site
$200K - $400K/yr
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit ...
Sr. Software Engineer - AI Triton Kernels
San Jose, CA · On-site
$180 - $300/hr
Triton is a widely adopted language and compiler for high-performance GPU kernels, powering major AI frameworks such as PyTorch, vLLM, and SGLang. As AI workloads increasingly rely on Triton-based ...
Sr. Software Engineer - AI Triton Kernels
San Jose, CA · On-site
$180 - $300/hr
Triton is a widely adopted language and compiler for high-performance GPU kernels, powering major AI frameworks such as PyTorch, vLLM, and SGLang. As AI workloads increasingly rely on Triton-based ...
Experience working on high-performance server systems--you'd be just as comfortable with the internals of VLLM as you would with a complex PyTorch codebase. • Significant performance engineering ...
Experience working on high-performance server systems--you'd be just as comfortable with the internals of VLLM as you would with a complex PyTorch codebase. • Significant performance engineering ...
Product Marketing Manager
San Francisco, CA · On-site
$181K/yr
Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM ...
Product Marketing Manager
San Francisco, CA · On-site
$181K/yr
Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM ...
Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving. * Experience with Kubernetes, Docker, Azure ML, Databricks ...
Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving. * Experience with Kubernetes, Docker, Azure ML, Databricks ...
Sr. Software Engineer - AI Triton Kernels
San Jose, CA · Hybrid
$143K - $189K/yr
Triton is a widely adopted language and compiler for high-performance GPU kernels, powering major AI frameworks such as PyTorch, vLLM, and SGLang. As AI workloads increasingly rely on Triton-based ...
Sr. Software Engineer - AI Triton Kernels
San Jose, CA · Hybrid
$143K - $189K/yr
Triton is a widely adopted language and compiler for high-performance GPU kernels, powering major AI frameworks such as PyTorch, vLLM, and SGLang. As AI workloads increasingly rely on Triton-based ...
Sr. Software Engineer - AI Triton Kernels
San Jose, CA · On-site
$127K/yr
Triton is a widely adopted language and compiler for high-performance GPU kernels, powering major AI frameworks such as PyTorch, vLLM, and SGLang. As AI workloads increasingly rely on Triton-based ...
Sr. Software Engineer - AI Triton Kernels
San Jose, CA · On-site
$127K/yr
Triton is a widely adopted language and compiler for high-performance GPU kernels, powering major AI frameworks such as PyTorch, vLLM, and SGLang. As AI workloads increasingly rely on Triton-based ...
Vllm information
What is a vLLM?
How does a vLLM engineer typically collaborate with data scientists and product teams during model deployment?
What are the key skills and qualifications needed to thrive as a machine learning engineer working with vLLM, and why are they important?
What is the difference between Vllm vs Data Analyst?
| Aspect | Vllm | Data Analyst |
|---|---|---|
| Required Credentials | Typically requires knowledge of machine learning, AI, and programming languages like Python or R | Requires skills in statistics, Excel, SQL, and data visualization tools |
| Work Environment | Often in tech companies, research labs, or AI-focused teams | Commonly in business, finance, healthcare, and marketing sectors |
| Industry Usage | Emerging role in AI and machine learning projects | Established role in data-driven decision making |
| Common Search/Comparison | Vllm vs Data Analyst |
The main difference between Vllm and Data Analyst lies in their focus and skill set. Vllm professionals specialize in AI and machine learning models, often working in tech environments, while Data Analysts focus on interpreting data to inform business decisions. Both roles require analytical skills, but Vllm roles demand programming and AI expertise, whereas Data Analysts emphasize statistical analysis and data visualization.
What are popular job titles related to Vllm jobs in California?
For Vllm jobs in California, the most frequently searched job titles are:
What job categories do people searching Vllm jobs in California look for?
The top searched job categories for Vllm jobs in California are:
What cities in California are hiring for Vllm jobs?
Cities in California with the most Vllm job openings:

Full-time
Re-posted 28 days ago
Adobe rating
8.9
Based on 10 frontline employees who took The Breakroom Quiz
40th of 246 rated software companies
Job description
Adobe Firefly's ASML group invites research scientists and engineers passionate about conditional generation and editing of large generative AI models. This role emphasizes images and videos. We strive to advance generative AI technology while guaranteeing models possess excellent quality and control.
As an Applied Scientist, you will define technical strategy for multimodal data intelligence systems, architect and optimize distributed LLM/VLM inference platforms, and develop innovative solutions for automated captioning, tagging, metadata enrichment, and dataset creation. You will work at the intersection of research and engineering, collaborating with teams across modeling, infrastructure, data, evaluation, and product to deliver high-quality AI capabilities at scale.
You will have the opportunity to influence the next generation of Adobe Firefly models by improving data quality, model efficiency, and scalable AI infrastructure used by millions of creators worldwide.
Job Responsibilities
- Architect and optimize distributed multimodal inference pipelines for large-scale image, video, and audio captioning, tagging, and metadata generation.
- Drive LLM/VLM inference optimization, including batching, scheduling, quantization, model serving, caching, and GPU utilization to maximize throughput and cost efficiency.
- Build scalable data generation workflows using state-of-the-art vision-language and multimodal foundation models to improve training data quality.
- Lead technical strategy for automated dataset annotation, filtering, quality scoring, deduplication, and metadata enrichment across multimodal datasets.
- Design distributed processing systems capable of handling billions of media assets across heterogeneous compute environments.
- Collaborate with research teams to productionize new LLM/VLM capabilities while ensuring scalability, reliability, and operational efficiency.
- Partner with infrastructure teams to improve distributed execution frameworks, storage systems, and inference services.
- Drive cross-functional alignment across data, research, infrastructure, evaluation, and product teams on multimodal data processing strategy.
- Mentor engineers in distributed systems, scalable ML infrastructure, and multimodal AI engineering best practices.
- Ph.D. or M.S. in Computer Science, Machine Learning, or a related technical field, with significant industry experience designing and deploying large-scale distributed ML systems.
- Deep expertise in large language models (LLMs), vision-language models (VLMs), or multimodal foundation models, with hands-on experience building, optimizing, and serving inference workloads at scale.
- Strong background in distributed systems, large-scale data processing, and cloud-native ML infrastructure, with experience using frameworks such as Ray, Spark, Dask, Kubernetes, or equivalent technologies.
- Proven experience optimizing large-scale LLM/VLM inference systems, including techniques such as batching, parallelism, quantization, model serving, GPU utilization optimization, and latency/throughput tuning.
- Experience building high-throughput multimodal data pipelines for automated image, video, and audio understanding tasks, including captioning, tagging, OCR, metadata extraction, and semantic indexing.
- Hands-on experience with modern ML inference and serving frameworks such as vLLM, TensorRT-LLM, SGLang, Triton Inference Server, TGI, or equivalent technologies.
- Experience managing and processing petabyte-scale multimodal datasets using distributed storage and data processing systems.
- Familiarity with multimodal embedding models, retrieval-augmented systems, vector search infrastructure, and data quality evaluation methodologies.
- Strong software engineering skills in Python and PyTorch, with a track record of developing reliable, production-grade distributed ML systems.
- Excellent communication and collaboration skills, with the ability to influence technical strategy and drive alignment across research, infrastructure, and product teams.
About Adobe
Adobe empowers everyone to create through innovative platforms and tools that unleash creativity, productivity and personalized customer experiences. Adobe's industry-leading offerings including Adobe Acrobat Studio, Adobe Express, Adobe Firefly, Creative Cloud, Adobe Experience Platform, Adobe Experience Manager, and GenStudio enable people and businesses to turn ideas into impact, powered by AI and driven by human ingenuity.
Our 30,000+ employees worldwide are creating the future and raising the bar as we drive the next decade of growth. We're on a mission to hire the very best and believe in creating a company culture where all employees are empowered to make an impact. At Adobe, we believe that great ideas can come from anywhere in the organization. The next big idea could be yours.
Let's Adobe together
At Adobe, we believe in creating a company culture where all employees are empowered to make an impact. Learn more about Adobe life, including our values and culture, focus on people, purpose and community, Adobe for All, comprehensive benefits programs, the stories we tell, the customers we serve, and how you can help us advance our mission of empowering everyone to create.
Adobe is proud to be an Equal Employment Opportunity employer. We do not discriminate based on gender, race or color, ethnicity or national origin, age, disability, religion, sexual orientation, gender identity or expression, veteran status, or any other protected characteristic. Learn more.
Adobe aims to make our Careers website and recruiting process accessible to any and all users. If you have a disability or special need that requires accommodation to navigate our website or complete the application process, email accommodations@adobe.com or call +1 408-536-3015.
AI Use Guidelines for Interviews:
Our interviews are designed to reflect your own skills and thinking. The use of AI or recording tools during live interviews is not permitted unless explicitly invited by the interviewer or approved in advance as part of a reasonable accommodation. If these tools are used inappropriately or in a way that misrepresents your work, your application may not move forward in the process.
At Adobe, we empower employees to innovate with AI - and we look for candidates eager to do the same. As part of the hiring experience, we provide clear guidance on where AI is encouraged during the process and where it's restricted during live interviews. See how we think about AI in the hiring experience.
Expected Pay Range:
Our compensation reflects the cost of labor across several U.S. geographic markets, and we pay differently based on those defined markets. The U.S. pay range for this position is $164,000 -- $313,300 annually. Pay within this range varies by work location and may also depend on job-related knowledge, skills, and experience. Your recruiter can share more about the specific salary range for the job location during the hiring process.
In California, the pay range for this position is $216,400 - $313,300
At Adobe, for sales roles starting salaries are expressed as total target compensation (TTC = base + commission), and short-term incentives are in the form of sales commission plans. Non-sales roles starting salaries are expressed as base salary and short-term incentives are in the form of the Annual Incentive Plan (AIP).
In addition, certain roles may be eligible for long-term incentives in the form of a new hire equity award.
State-Specific Notices:
California:
Fair Chance Ordinances
Adobe will consider qualified applicants with arrest or conviction records for employment in accordance with state and local laws and "fair chance" ordinances.
Colorado:
Application Window Notice
If this role is open to hiring in Colorado (as listed on the job posting), the application window will remain open until at least the date and time stated above in Pacific Time, in compliance with Colorado pay transparency regulations. If this role does not have Colorado listed as a hiring location, no specific application window applies, and the posting may close at any time based on hiring needs.
Massachusetts:
Massachusetts Legal Notice
It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.
About Adobe
Sourced by ZipRecruiter
Adobe for All is our vision to advance diversity, equity, and inclusion (DEI) across our company and in our communities. We’re focused on creating a more diverse and inclusive workforce; unleashing the full potential of every employee; and driving meaningful impact for Adobe, our industry, and society at large. Creativity has the power to unite us and inspire us to change the world. Through a vision we call Creativity for All, we’re empowering millions of people of all ages and backgrounds to express themselves, reach their full potential, and share their diverse perspectives with the world. We’re committed to advancing the responsible use of technology and driving a positive environmental impact through sustainability and climate action. Our innovations are making a significant impact across AI ethics, security, privacy, trust and safety, accessibility, and sustainability.
Industry
Computer and computer peripheral equipment and software wholesalers
Company size
10,000+ Employees
Headquarters location
San Jose, CA, US
Year founded
1982