1

Vllm Jobs in Texas (NOW HIRING)

Senior AI Engineer Agentic Systems

Plano, TX · On-site

$100K - $137K/yr

... vLLM alongside commercial LLM APIs. • Implement solutions for regulated, high-trust environments, including PII handling, auditability, deterministic fallbacks, and observability. Required ...

AI Infrastructure Engineer

Austin, TX · On-site

$170K - $315K/yr

Upstream your architectural improvements and hardware backends directly into open-source repositories like vLLM, SGLang, and PyTorch, acting as a bridge between the hardware teams and the open-source ...

Optimize inference using ONNX, vLLM, TensorRT, and GPU acceleration. * Manage datasets, preprocess data, and implement RAG with vector databases (FAISS, Chroma, Pinecone). * Automate training ...

Integrating networking capabilities into AI serving stacks such as vLLM, SGLang, and TensorRT-LLM. * Publishing findings, representing NVIDIA in industry forums and standards bodies, and mentoring ...

AI Engineer

Fort Worth, TX · On-site

$150 - $230/hr

Deploy and optimize LLM / VLLM inference stacks using vLLM, Hugging Face Transformers, TensorRT-LLM, TGI, Ollama, llama.cpp, Ray Serve, or Triton Inference Server. * Integrate commercial and ...

AI Engineer

Fort Worth, TX · On-site

$114K - $150K/yr

Deploy and optimize LLM / VLLM inference stacks using vLLM, Hugging Face Transformers, TensorRT-LLM, TGI, Ollama, llama.cpp, Ray Serve, or Triton Inference Server. * Integrate commercial and open ...

LLM Infrastructure Engineer

Houston, TX · On-site

$97K - $127K/yr

Experience with OpenAI-compatible APIs, vLLM, or Text Generation Inference (TGI) * Experience deploying AI workloads on Azure GPU infrastructure * Familiarity with LoRA / PEFT fine-tuning * Exposure ...

Sr. ML Engineer

Austin, TX · Hybrid

$130K/yr

Build robust serving infrastructure to productionize machine learning models and Large Language Models (LLMs) using modern serving frameworks (e.g., vLLM, TensorRT-LLM, KServe, Triton). * Cloud ...

next page

Showing results 1-20

Vllm information

What is a vLLM?

VLLM stands for 'Virtual Large Language Model.' In the context of AI development, VLLM professionals work with optimized inference engines for large language models, enabling faster and more efficient deployment of AI models in production environments. Their responsibilities often include integrating LLMs into applications, optimizing model performance, and ensuring scalability for real-time use cases. They may also collaborate with data scientists and engineers to manage resources and streamline AI workflows.

How does a vLLM engineer typically collaborate with data scientists and product teams during model deployment?

VLLM Engineers work closely with data scientists to understand the specific requirements and fine-tuning needs of large-scale language models. They are often responsible for integrating these models into production systems, ensuring scalability and efficiency. Collaboration with product teams is crucial to align model capabilities with user needs and to troubleshoot real-world application challenges. Frequent communication and agile workflows are common, as updates or optimizations may be needed rapidly based on feedback from both teams.

What are the key skills and qualifications needed to thrive as a machine learning engineer working with vLLM, and why are they important?

To thrive as a Machine Learning Engineer specializing in vLLM (a high-throughput LLM inference library), you need a strong understanding of machine learning principles, deep learning frameworks, and experience with Python programming. Familiarity with tools like PyTorch, CUDA, distributed computing, and cloud platforms, as well as relevant certifications in ML or data engineering, is highly valuable. Strong problem-solving, collaboration, and communication skills are essential for optimizing model performance and integrating with cross-functional teams. These capabilities ensure effective deployment and scaling of large language models, driving innovation and efficiency in AI applications.

What is the difference between Vllm vs Data Analyst?

AspectVllmData Analyst
Required CredentialsTypically requires knowledge of machine learning, AI, and programming languages like Python or RRequires skills in statistics, Excel, SQL, and data visualization tools
Work EnvironmentOften in tech companies, research labs, or AI-focused teamsCommonly in business, finance, healthcare, and marketing sectors
Industry UsageEmerging role in AI and machine learning projectsEstablished role in data-driven decision making
Common Search/ComparisonVllm vs Data Analyst

The main difference between Vllm and Data Analyst lies in their focus and skill set. Vllm professionals specialize in AI and machine learning models, often working in tech environments, while Data Analysts focus on interpreting data to inform business decisions. Both roles require analytical skills, but Vllm roles demand programming and AI expertise, whereas Data Analysts emphasize statistical analysis and data visualization.

What are popular job titles related to Vllm jobs in Texas?

For Vllm jobs in Texas, the most frequently searched job titles are:

What cities in Texas are hiring for Vllm jobs?

Cities in Texas with the most Vllm job openings:

Infographic showing various Vllm job openings in Texas as of August 2026, with employment types broken down into 95% Full Time, 3% Part Time, and 2% Contract. Highlights an 77% Physical, 6% Hybrid, and 17% Remote job distribution.

Senior AI Engineer Agentic Systems

Compunnel

Plano, TX • On-site

$100K - $137K/yr

Contractor

Posted 4 days ago


Job description

Job Summary
We are seeking a Senior AI Engineer to design and build production-grade agentic systems for customer-facing use cases. This role will focus on developing reliable AI agents capable of multi-turn conversation, tool calling, retrieval, escalation, and human handoff within a regulated, high-trust environment. The ideal candidate will have proven experience shipping LLM-based agentic systems to production and strong expertise across Python, Java, LangChain, LangGraph, RAG, model serving, evaluation, and distributed systems.
Key Responsibilities
• Design and build production agentic workflows for customer-facing use cases, including multi-turn conversation, tool calling, retrieval, escalation, and human handoff.
• Own agent orchestration end to end using LangGraph, LangChain, or equivalent frameworks, including state management, guardrails, and failure recovery.
• Integrate agents with core service layers, including Java/Spring microservices, APIs, and event streams.
• Build and tune evaluation loops, including automated evaluations, regression suites, trace analysis, and A/B measurement of agent quality.
• Work on inference and serving layers, including model routing, prompt and context management, and self-hosted serving with vLLM alongside commercial LLM APIs.
• Implement solutions for regulated, high-trust environments, including PII handling, auditability, deterministic fallbacks, and observability.
Required Qualifications
• 7+ years of software engineering experience, with senior or lead-level ownership of production systems.
• Proven hands-on experience delivering LLM-based agentic systems in production, including the ability to explain architecture, failure modes, and evaluation strategies for shipped systems.
• Strong Python experience as a primary agent and ML development language and Java experience for service integration.
• Advanced experience with LangChain, LangGraph, prompt and context engineering, tool/function calling, and RAG pipelines.
• Working knowledge of PyTorch and model fundamentals, including the ability to fine-tune models, debug model behavior, and evaluate technical tradeoffs.
• Experience with LLM serving and inference using vLLM or comparable technologies such as TGI or TensorRT-LLM.
• Experience working with commercial LLM APIs such as OpenAI, Anthropic, Bedrock, or Vertex.
• Strong distributed systems fundamentals, including APIs, queues, caching, observability, CI/CD, and production engineering practices.
• Advanced experience with GenAI prompt engineering.
• Advanced experience with OpenAI engineering.
• Advanced experience with testing, DevOps, and CI/CD practices.
Preferred Qualifications
• Hands-on experience with agent-builder platforms such as Sierra or Decagon, including deploying, extending, or evaluating platforms for customer-service use cases.
• Experience with voice agents, real-time or streaming inference, or contact-center integrations.
• Experience working in regulated industries such as finance or healthcare, including model risk, compliance review, and audit trails.
• Experience with evaluation frameworks such as LangSmith, Braintrust, or custom evaluation harnesses.
• Experience with LLM safety and guardrail tooling.

Compunnel logo

About Compunnel

Sourced by ZipRecruiter

Compunnel is a well-known company located in Plainsboro, NJ, US, recognized in the industry of IT Services and Solutions. Established in 1989, Compunnel offers a suite of services that help businesses integrate technology efficiently into their operations, a recognizable name in the IT solutions sphere for over three decades. The company’s service portfolio includes Digital Transformation, Business Intelligence, Cloud Services, Cybersecurity, and Application Modern Services, among others. Guided by its mission "to innovate with industry-leading digital solutions and disruptive tech strategies for unimagining business growth," the company underlines its commitment to offering out-of-the-box solutions to its clients. Remarkable achievements of the company include serving more than 30 Fortune 500 companies and providing job opportunities for over 50,000 individuals.

Industry

It services

Company size

501 - 1,000 Employees

Headquarters location

Plainsboro, NJ, US

Year founded

1994

Social media