1

Vllm Jobs in Utah (NOW HIRING)

Senior Machine Learning Engineer

Sandy, UT · Hybrid

$99K - $136K/yr

LoRA/PEFT for speech models, inference optimization (quantization, SGLang/vLLM serving for audio, distillation), experience with at least one open-source TTS family * GPU cost modeling * Proficiency ...

Senior Machine Learning Engineer

Lehi, UT · On-site +1

$98K - $134K/yr

Familiarity with frameworks such as vLLM, DeepSpeed, FSDP, or similar technologies. * Experience working with enterprise or domain-specific AI applications. * Bachelor's or advanced degree in ...

Senior Machine Learning Engineer

Lehi, UT · On-site +1

$144K - $233K/yr

Familiarity with frameworks such as vLLM, DeepSpeed, FSDP, or similar technologies. * Experience working with enterprise or domain-specific AI applications. * Bachelor's or advanced degree in ...

Senior Machine Learning Engineer

Lehi, UT · On-site

$144K - $233K/yr

Familiarity with frameworks such as vLLM, DeepSpeed, FSDP, or similar technologies. * Experience working with enterprise or domain-specific AI applications. * Bachelor's or advanced degree in ...

Vllm information

What is a vLLM?

VLLM stands for 'Virtual Large Language Model.' In the context of AI development, VLLM professionals work with optimized inference engines for large language models, enabling faster and more efficient deployment of AI models in production environments. Their responsibilities often include integrating LLMs into applications, optimizing model performance, and ensuring scalability for real-time use cases. They may also collaborate with data scientists and engineers to manage resources and streamline AI workflows.

How does a vLLM engineer typically collaborate with data scientists and product teams during model deployment?

VLLM Engineers work closely with data scientists to understand the specific requirements and fine-tuning needs of large-scale language models. They are often responsible for integrating these models into production systems, ensuring scalability and efficiency. Collaboration with product teams is crucial to align model capabilities with user needs and to troubleshoot real-world application challenges. Frequent communication and agile workflows are common, as updates or optimizations may be needed rapidly based on feedback from both teams.

What are the key skills and qualifications needed to thrive as a machine learning engineer working with vLLM, and why are they important?

To thrive as a Machine Learning Engineer specializing in vLLM (a high-throughput LLM inference library), you need a strong understanding of machine learning principles, deep learning frameworks, and experience with Python programming. Familiarity with tools like PyTorch, CUDA, distributed computing, and cloud platforms, as well as relevant certifications in ML or data engineering, is highly valuable. Strong problem-solving, collaboration, and communication skills are essential for optimizing model performance and integrating with cross-functional teams. These capabilities ensure effective deployment and scaling of large language models, driving innovation and efficiency in AI applications.

What is the difference between Vllm vs Data Analyst?

AspectVllmData Analyst
Required CredentialsTypically requires knowledge of machine learning, AI, and programming languages like Python or RRequires skills in statistics, Excel, SQL, and data visualization tools
Work EnvironmentOften in tech companies, research labs, or AI-focused teamsCommonly in business, finance, healthcare, and marketing sectors
Industry UsageEmerging role in AI and machine learning projectsEstablished role in data-driven decision making
Common Search/ComparisonVllm vs Data Analyst

The main difference between Vllm and Data Analyst lies in their focus and skill set. Vllm professionals specialize in AI and machine learning models, often working in tech environments, while Data Analysts focus on interpreting data to inform business decisions. Both roles require analytical skills, but Vllm roles demand programming and AI expertise, whereas Data Analysts emphasize statistical analysis and data visualization.

What are popular job titles related to Vllm jobs in Utah?

For Vllm jobs in Utah, the most frequently searched job titles are:

What cities in Utah are hiring for Vllm jobs?

Cities in Utah with the most Vllm job openings:

Internship - Machine Learning Engineer

Salt Lake City, UT • On-site

Smule
11 - 50 employees

$110 - $150/hr

Other

Re-posted 29 days ago


Job description

Smule has been on a mission to bring the world together through music since 2008. Music is much more than listening… it's about creating, sharing, discovering, participating, and connecting with people. With dozens of millions of monthly active users creating over 20 million songs every day, Smule is connecting people all over the world through the joy of making music and transforming the music landscape from one of passive listening to collaborative creative expression and active engagement.

About the Role:

We are looking for a Machine Learning Engineer to own the end-to-end lifecycle of ML models in production at Smule, from training and optimization through deployment, monitoring, and iteration. You will work closely with research scientists to bring models off the bench and into scalable, reliable systems that serve millions of users. The ideal candidate is a strong engineer first, with deep practical knowledge of ML systems, a passion for reliability, and an eye for performance.

We strongly encourage candidates with non-traditional backgrounds to apply. If your path into ML engineering came through backend systems, DevOps, audio software, data engineering, or another field, we want to hear from you.

What You'll Be Doing:
  • Design, build, and maintain production ML pipelines encompassing data ingestion, feature engineering, model training, evaluation, and deployment.
  • Optimize models for production constraints including latency, throughput, memory footprint, and cost, using techniques such as quantization, distillation, pruning, and efficient serving architectures.
  • Implement robust monitoring, alerting, and observability for deployed models, covering data drift, prediction quality, and system health.
  • Collaborate with research scientists to integrate new model architectures and training techniques into production systems with minimal friction.
  • Build and improve CI/CD pipelines for ML, including automated testing, validation gates, and staged rollouts.
  • Manage compute infrastructure and costs, making informed tradeoffs between performance, reliability, and budget.
What We're Looking For:
  • Degree (B.S., M.S., or Ph.D.) in Computer Science, Software Engineering, Electrical Engineering, or a related technical discipline, or currently pursuing one.
  • Strong proficiency in Python and experience with deep learning serving (TorchServe, Triton, vLLM, or equivalent).
  • Solid understanding of systems engineering: networking, storage, containerization, orchestration, and monitoring.
  • Ability to reason about tradeoffs between latency, throughput, cost, and model quality.
Bonus Points For:
  • Experience serving large language models or other generative models at scale.
  • Familiarity with audio/music processing pipelines and real-time inference constraints.
  • Experience with Bayesian optimization, bandit algorithms, or adaptive experimentation platforms.
  • Contributions to open-source ML infrastructure projects.

Smule is an Equal Opportunity Employer and considers all qualified applicants without regard to race, color, religion, sex, gender identity or expression, sexual orientation, national origin, ancestry, age, disability, medical condition, genetic information, marital status, military or veteran status, or any other protected characteristic under federal, state, or local law.

We are committed to creating an inclusive environment for all employees and applicants. If you require a reasonable accommodation during the application or interview process, please let us know.

#J-18808-Ljbffr