Senior AI Engineer
$100K - $138K/yr
Engineer and tune LLM inference serving stacks -- primary depth in vLLM with breadth across the inference ecosystem -- for client latency, throughput, and cost targets. * Tune inference performance ...
Quick apply
$100K - $138K/yr
Engineer and tune LLM inference serving stacks -- primary depth in vLLM with breadth across the inference ecosystem -- for client latency, throughput, and cost targets. * Tune inference performance ...
Quick apply
$100K - $138K/yr
Engineer and tune LLM inference serving stacks -- primary depth in vLLM with breadth across the inference ecosystem -- for client latency, throughput, and cost targets. * Tune inference performance ...
Atlanta, GA · On-site
$48.25 - $66.50/hr
GPU compute (NVIDIA CUDA/DCGM, vllm, llm-d, Intel Level-Zero/ driver stack and SR-IOV)
Atlanta, GA · On-site
$48.25 - $66.50/hr
GPU compute (NVIDIA CUDA/DCGM, vllm, llm-d, Intel Level-Zero/ driver stack and SR-IOV)
Atlanta, GA · On-site
GPU compute (NVIDIA CUDA/DCGM, vllm, llm-d, Intel Level-Zero/ driver stack and SR-IOV)
Atlanta, GA · On-site
GPU compute (NVIDIA CUDA/DCGM, vllm, llm-d, Intel Level-Zero/ driver stack and SR-IOV)
Alpharetta, GA · On-site
$55.75 - $74/hr
Deploy, scale, and manage LLM inference servers (e.g., vLLM, Ray Serve, NVIDIA Triton) on Kubernetes across multi-cloud environments. * Implement comprehensive observability, logging, and tracing for ...
Alpharetta, GA · On-site
$55.75 - $74/hr
Deploy, scale, and manage LLM inference servers (e.g., vLLM, Ray Serve, NVIDIA Triton) on Kubernetes across multi-cloud environments. * Implement comprehensive observability, logging, and tracing for ...
Alpharetta, GA · On-site
$105K - $137K/yr
Deploy and manage containerized AI applications and model inference servers (e.g., vLLM, Ray Serve, NVIDIA Triton) on Kubernetes across multi-cloud environments (AWS, GCP, Azure). * Implement ...
Alpharetta, GA · On-site
$105K - $137K/yr
Deploy and manage containerized AI applications and model inference servers (e.g., vLLM, Ray Serve, NVIDIA Triton) on Kubernetes across multi-cloud environments (AWS, GCP, Azure). * Implement ...
Atlanta, GA · On-site
$140 - $210/hr
Experienced with LLM inference stacks (TensorRT-LLM, VLLM, llama.cpp) and developing RAG systems. * Strong problem-solving and analytical skills, with excellent communication to foster cross ...
New
Atlanta, GA · On-site
$140 - $210/hr
Experienced with LLM inference stacks (TensorRT-LLM, VLLM, llama.cpp) and developing RAG systems. * Strong problem-solving and analytical skills, with excellent communication to foster cross ...
New
Atlanta, GA · On-site
$140 - $190/hr
Work with production‑ready inference runtimes such as vLLM, ONNX Runtime, and NVIDIA Triton. * Contribute to model conversion, quantization, and optimization for efficient inference. * Partner with ...
Atlanta, GA · On-site
$140 - $190/hr
Work with production‑ready inference runtimes such as vLLM, ONNX Runtime, and NVIDIA Triton. * Contribute to model conversion, quantization, and optimization for efficient inference. * Partner with ...
Atlanta, GA · On-site
$103K - $135K/yr
Architect and deploy with NVIDIA platform tools including Base Command Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM ...
Atlanta, GA · On-site
$103K - $135K/yr
Architect and deploy with NVIDIA platform tools including Base Command Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM ...
Atlanta, GA · On-site
... vLLM, TensorRT-LLM, or TGI. Small language models & open-weight models • Train and optimize open-weight models such as Llama, Qwen, Mistral, or DeepSeek; build specialized small language models ...
Atlanta, GA · On-site
... vLLM, TensorRT-LLM, or TGI. Small language models & open-weight models • Train and optimize open-weight models such as Llama, Qwen, Mistral, or DeepSeek; build specialized small language models ...
Atlanta, GA · On-site
$100K - $138K/yr
Work with production-ready inference runtimes such as vLLM, ONNX Runtime, and NVIDIA Triton. * Contribute to model conversion, quantization, and optimization for efficient inference. * Partner with ...
Atlanta, GA · On-site
$100K - $138K/yr
Work with production-ready inference runtimes such as vLLM, ONNX Runtime, and NVIDIA Triton. * Contribute to model conversion, quantization, and optimization for efficient inference. * Partner with ...
Atlanta, GA · On-site
Optimize inference performance - latency, throughput, quantization, and deployment efficiency - for production, including frameworks such as vLLM, TensorRT-LLM, or TGI. Small language models & open ...
Atlanta, GA · On-site
Optimize inference performance - latency, throughput, quantization, and deployment efficiency - for production, including frameworks such as vLLM, TensorRT-LLM, or TGI. Small language models & open ...
Atlanta, GA · On-site
$60.50 - $79.75/hr
Familiarity one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
Quick apply
Atlanta, GA · On-site
$60.50 - $79.75/hr
Familiarity one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
Atlanta, GA · On-site
$60.50 - $79.75/hr
Familiarity one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
Atlanta, GA · On-site
$60.50 - $79.75/hr
Familiarity one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
Atlanta, GA · On-site
$182 - $242/hr
Familiarity one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
Atlanta, GA · On-site
$182 - $242/hr
Familiarity one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
Atlanta, GA · On-site
$182 - $242/hr
Familiarity with one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
Atlanta, GA · On-site
$182 - $242/hr
Familiarity with one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
Atlanta, GA · On-site
$182K - $242K/yr
Familiarity one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
Quick apply
Atlanta, GA · On-site
$182K - $242K/yr
Familiarity one or more deep learning frameworks (PyTorch) and modern LLM stack (VLLM, langchain / LlamaIndex) * Experience using Slurm or Kubernetes for ML job orchestration * Experience with ...
| Aspect | Vllm | Data Analyst |
|---|---|---|
| Required Credentials | Typically requires knowledge of machine learning, AI, and programming languages like Python or R | Requires skills in statistics, Excel, SQL, and data visualization tools |
| Work Environment | Often in tech companies, research labs, or AI-focused teams | Commonly in business, finance, healthcare, and marketing sectors |
| Industry Usage | Emerging role in AI and machine learning projects | Established role in data-driven decision making |
| Common Search/Comparison | Vllm vs Data Analyst |
The main difference between Vllm and Data Analyst lies in their focus and skill set. Vllm professionals specialize in AI and machine learning models, often working in tech environments, while Data Analysts focus on interpreting data to inform business decisions. Both roles require analytical skills, but Vllm roles demand programming and AI expertise, whereas Data Analysts emphasize statistical analysis and data visualization.
For Vllm jobs in Georgia, the most frequently searched job titles are:
The top searched job categories for Vllm jobs in Georgia are:
Cities in Georgia with the most Vllm job openings:

At BlueAlly, our mission is to make technology more accessible, more certain, and more impactful for every organization.
From cloud to cybersecurity, infrastructure to application modernization, we thrive on cutting-edge technologies and services. Elevate the impact of technology across your enterprise with world-class expertise that produces game-changing insights. Turn complex decisions into clear opportunities with a trusted guide to technology that ensures the next digital advance will be your decisive advantage. Trade IT complexity for capability with solutions that elevate possibilities, and advance with certainty, knowing you have BlueAlly as your ally in next. BlueAlly. Conquer Complexity.
Job DescriptionWe are hiring a Senior AI Engineer to design, build, and operate enterprise AI systems across our client portfolio. You will work end-to-end across the AI stack — from inference engines and platform infrastructure (vLLM, KV cache, Dynamo-style serving, GPU-accelerated AI Factory platforms) up through application-level engineering (RAG pipelines, agent workflows, prompt engineering, evaluation methodology).
This role is for an engineer who can lead workstreams independently, mentor more junior engineers, and serve as the technical authority that clients trust to deliver production AI outcomes. You'll engage directly with client architects, data scientists, application teams, and executives — and you'll leave each engagement having raised both the client's capability and BlueAlly's practice.
Key Responsibilities:
Preferred Qualifications:
Certifications (Preferred):
What Sets You Apart:
About BlueAlly
BlueAlly is a leading provider of IT services and solutions, helping organizations conquer IT complexity across cloud, cybersecurity, infrastructure, data, and application modernization. Headquartered in Atlanta,Georgia, with delivery teams across the United States and globally, BlueAlly serves clients ranging from mid-market enterprises to large public-sector and commercial organizations.
Founded in 2011, BlueAlly delivers across the full technology lifecycle — from strategy and design through implementation, managed services, and continuous optimization. The company is recognized on CRN's Tech Elite 150 and MSP 500 lists and partners deeply with leading technology vendors. As enterprise AI moves from pilot to production, BlueAlly is investing in the people, platforms, and practices required to deliver AI Factory outcomes for our clients — and this role is at the center of that investment.
Equal Employment Opportunity
BlueAlly is an Equal Opportunity Employer. We are committed to building a diverse and inclusive workforce and to making employment decisions based on merit, qualifications, and business need. BlueAlly does not discriminate in employment on the basis of race, color, religion, sex (including pregnancy), national origin, age, disability, genetic information, sexual orientation, gender identity or expression, marital status, veteran status, or any other protected characteristic under applicable federal, state, or local law.
BlueAlly provides reasonable accommodations to qualified applicants and employees with disabilities. If you require an accommodation to participate in the application or interview process, please contact our People team.