Charlotte, United States | Posted on 07/15/2026
Dallas, TX or Charlotte, NC or Raleigh, NC
Role Overview
We are seeking a highly skilled Generative AI Engineer with a strong Python background to design, develop, and deploy cutting‑edge AI solutions. The ideal candidate will have hands‑on experience with Large Language Models (LLMs), Vision Language Models (Vision LLMs/VLMs), vLLM inference framework, prompt engineering, and modern Generative AI frameworks, along with proven expertise in building scalable AI applications for enterprise use cases.
This role focuses on developing Agentic AI systems, Retrieval‑Augmented Generation (RAG), multimodal AI solutions, and high‑performance LLM inference while integrating GenAI capabilities into production‑grade enterprise applications.
Mission
Design and deliver scalable, production‑ready Generative AI solutions leveraging modern LLMs, Vision LLMs, Agentic AI frameworks, RAG architectures, and cloud AI platforms to power intelligent enterprise applications.
Key Responsibilities Design and implement Generative AI solutions for:
- Text‑based AI applications
- Image‑based AI applications
- Multimodal AI applications
- Develop and optimize advanced prompt engineering strategies to improve LLM performance, accuracy, and reliability.
- Build and integrate embedding‑based retrieval systems and Retrieval‑Augmented Generation (RAG) pipelines.
Design and implement
Agentic AI applications including:
- Context management
- Session and memory handling
- Tool calling and workflow orchestration
- Deploy and optimize vLLM for high‑throughput, low‑latency LLM inference in production environments.
- Build scalable APIs using Python and integrate GenAI capabilities into enterprise applications and workflows.
- Collaborate with cross‑functional teams to deploy AI solutions at scale.
- Ensure AI solutions are secure, scalable, reliable, and production‑ready.
Required Qualifications Programming
- Strong proficiency in Python
Solid experience with AI/ML frameworks including:
Agentic AI
Hands‑on experience building multi‑agent AI systems, including:
- Session management
- Memory handling
- Tool integration and orchestration
Practical experience with:
- vLLM for optimized LLM serving and inference
Retrieval & Search
Experience with:
- Embeddings
- Retrieval‑Augmented Generation (RAG)
- Semantic Search
Experience with one or more:
MLOps
- Understanding of MLOps and LLMOps practices
- Experience deploying scalable AI applications in production
Preferred Qualifications
- Experience with multimodal AI systems combining text, images, and documents
Knowledge of AI ethics, including:
- Responsible AI practices
- Experience designing AI systems with governance, transparency, and compliance in mind
- Experience with distributed GPU inference, model optimization, quantization, and high‑performance AI serving
- Familiarity with frameworks such as LangChain, LangGraph, LlamaIndex, CrewAI, or AutoGen