1

Rag Llm Jobs (NOW HIRING)

Senior Data Engineer, AI Platform

San Jose, CA · On-site

$124K - $168K/yr

Build and optimize retrieval pipelines for RAG and LLM-based applications * Design and manage vector data pipelines (embedding generation, indexing, storage, retrieval) * Implement hybrid retrieval ...

AI/ML Engineer

Plano, TX · On-site

$109K - $131K/yr

This role offers the opportunity to work on cutting-edge AI initiatives, including Generative AI, Retrieval-Augmented Generation (RAG), LLM fine-tuning, and intelligent enterprise applications. Key ...

AI/LLM Eng 12+ Months New jersey Must have experience in: Large Language Models (OpenAI/Azure ... RAG architectures Agentic AI frameworks (LangChain, LangGraph, CrewAI, AutoGen or similar) Vector ...

Showing results 41-60

Rag Llm information

See salary details

$45K

$75.3K

$110K

How much do rag llm jobs pay per year?

As of Aug 8, 2026, the average yearly pay for rag llm in the United States is $75,300.00, according to ZipRecruiter salary data. Most workers in this role earn between $62,000.00 and $87,000.00 per year, depending on experience, location, and employer.

What is the difference between Rag Llm vs Data Scientist?

AspectRag LlmData Scientist
Required CredentialsTypically a master's or PhD in AI, machine learning, or related fieldsUsually a master's or PhD in data science, statistics, or computer science
Work EnvironmentResearch labs, AI development teams, tech companiesBusiness analytics, research, tech firms, consulting
Industry UsageAI research, natural language processing, machine learning projectsData analysis, predictive modeling, data-driven decision making

Rag Llm and Data Scientist roles often overlap in AI and data analysis fields, but Rag Llm focuses more on language models and AI research, while Data Scientists handle broader data analysis and modeling tasks. Both require advanced degrees and work in tech-driven environments, but their core responsibilities differ in scope and application.

What is a RAG LLM?

RAG LLMs, or Retrieval-Augmented Generation Large Language Models, are advanced AI systems that combine the strengths of traditional language models with external data retrieval systems. They work by first searching a relevant database or knowledge base for up-to-date information, and then using a language model to generate responses based on both the retrieved content and their own training. This approach helps LLMs provide more accurate, current, and contextually relevant answers, especially for specialized or rapidly changing topics. RAG LLMs are widely used in customer support, research, and enterprise applications to improve information accuracy and reliability.

How do RAG LLM engineers collaborate with data scientists and product teams to improve retrieval-augmented generation systems?

RAG LLM engineers often work closely with data scientists to fine-tune retrieval mechanisms, optimize model performance, and evaluate system outputs. They also collaborate with product teams to understand user needs, integrate feedback, and ensure the system delivers relevant, accurate information. Regular cross-functional meetings and code reviews are common, fostering a collaborative environment focused on continuous improvement and innovation in response to real-world challenges.

What are the key skills and qualifications needed to thrive as a Retrieval-Augmented Generation (RAG) LLM engineer?

To thrive as a Retrieval-Augmented Generation (RAG) LLM Engineer, you need a strong background in natural language processing, machine learning, and software development, often supported by a degree in computer science or a related field. Familiarity with frameworks like PyTorch, Hugging Face Transformers, vector databases, and cloud platforms, along with experience deploying large language models, is essential. Analytical thinking, problem-solving abilities, and effective communication are crucial soft skills for collaboration and innovation in this fast-evolving space. These skills ensure the development of robust, scalable, and accurate retrieval-augmented AI systems that meet real-world information needs.
More about Rag Llm jobs
What cities are hiring for Rag Llm jobs? Cities with the most Rag Llm job openings:
What states have the most Rag Llm jobs? States with the most job openings for Rag Llm jobs include:
Infographic showing various Rag Llm job openings in the United States as of August 2026, with employment types broken down into 95% Full Time, 1% Part Time, and 4% Contract. Highlights an 76% Physical, 5% Hybrid, and 19% Remote job distribution, with an average salary of $75,300 per year, or $36.2 per hour.

Technology Architect | Cloud Platform | Google Cloud - Architecture

Spruce Infotech

Charlotte, NC • On-site

$63.50 - $80.75/hr

Full-time

Posted 25 days ago


Job description

POC: Sam Chavez
ATTENTION ALL SUPPLIERS!!!
READ BEFORE SUBMITTING:
• UPDATED CONTACT NUMBER and EMAIL ID is a MANDATORY REQUEST from our client for all the submissions
• We prioritize endorsing those with complete and accurate information
• Avoid submitting duplicate profiles. We will Reject/Disqualify immediately.
• Make sure that candidate's interview schedules are updated. Please inform the candidate to keep their lines open.
• Please submit profiles within the max proposed rate.
• Please make sure to TAG the profiles correctly if the candidate has WORKED FOR INFOSYS as a SUBCON or FTE.
MANDATORY: Please include in the resume the candidate's complete & updated contact information (Phone number, Email address and Skype ID) as well as a set of 5 interview timeslots over a 72-hour period after submitting the profile when the hiring managers could potentially reach to them. PROFILES WITHOUT THE REQUIRED DETAILS and TIME SLOTS will be REJECTED.
Job Title: Technology Architect | Cloud Platform | Google Cloud - Architecture - Gen AI Engineer
Work Location & Reporting Address: Charlotte, NC 28202 (Onsite-Hybrid. LOCAL CANDIDATES ONLY!!!)
Contract duration: 12
MAX VENDOR RATE: XXX per hour max
Target Start Date: 01 Jul 2026
Does this position require Visa independent candidates only? Yes
Must Have Skills:
• GEN AI
• Agentic AI
• VLLM
• fAST API
• REST API
• MCD
• Lang Graph
• Lang Chain
• Graph RAG
• ML Ops
• Python
• ML
• Data Science
• RAG
• LLM
Nice to Have Skills:
• GCP
• Prompt Engineering
Detailed Job Description:
We are seeking a highly skilled Generative AI Engineer with a strong Python background to design, develop, and deploy cutting-edge AI solutions. The ideal candidate will have hands-on experience with Large Language Models (LLMs), prompt engineering, and Gen AI frameworks, along with expertise in building scalable AI applications. Experience in Developing Agentic AI solutions.
Key Responsibilities:
• Design and implement Generative AI models for text, image, or multimodal applications.
• Develop prompt engineering strategies and embedding-based retrieval systems.
• Integrate Gen AI capabilities into web applications and enterprise workflows.
• Build agentic AI applications with context engineering and MCP tools.
Required Skills & Qualifications:
• 7+ years of hands-on experience in AI, Data science, ML, GEN AI
• 2 years of strong hands on experience in Agentic AI, VLLM's, GEN AI, Lang Chain, Lang Graph, RAG, LLM OPS and AI Services in GCP and Azure.
• Strong hands on experience designing and deploying Retrieval-Augmented Generation (RAG) pipelines
• Strong MLOps/LLMOps experience with CI/CD automation,
• Extensive experience with LangChain, LangGraph, and agentic AI patterns including routing, memory, multi-agent orchestration, guardrails, and failure recovery.
• Experience in Cloud-native engineering across AWS (SageMaker, Lambda, ECS/Fargate, S3, API Gateway, Step Functions) and GCP (Vertex AI) for scalable AI delivery
• Experience in Developing microservices and API development using FastAPI, REST APIs, Pydantic/JSON schemas, Docker, and Kubernetes for low-latency serving.
• Strong Hands-on experience with vector databases and semantic search technologies including Pinecone, FAISS, ChromaDB, and embedding lifecycle management
• Strong proficiency in Python and AI/ML frameworks (PyTorch, TensorFlow).
• Hands on experience using session and memory for building multi-agent systems along with using MCP tools.
• Hands-on experience with LLMs, transformers, and Hugging Face ecosystem.
• Knowledge and experience with vector databases and RAG technique for semantic search.
• Familiarity with cloud AI services (AWS SageMaker, Azure OpenAI, GCP Vertex AI).
• Understanding of MLOps practices for scalable AI deployment.
• Strong experience in working with LLM fine-tuning with LoRA, QLoRA, PEFT,
• Strong experience in Architected advanced RAG systems using Pinecone, FAISS, Weaviate, Chroma, hybrid retrieval, and custom embeddings,
• Strong experience in Designing end-to-end LLMOps/MLOps pipelines using MLflow, DVC, SageMaker Pipelines, Vertex AI Pipelines, and GitHub Actions
• Experience in using cloud-native AI systems on AWS (SageMaker, Lambda, EKS, EC2, Step Functions, S3, Glue) and GCP Vertex AI, supporting high-volume inference and secure enterprise operations
• Experience in developing multi-agent orchestration workflows using LangGraph and CrewAI for tool-calling, validation agents, automated reasoning, and workflow supervision
Minimum Years of Experience:
• 10+ years
Certifications Needed:
Top 3 responsibilities you would expect the Subcon to shoulder and execute:
• Strong experience in GEN AI, LLM, RAG,ML, DL,ML Ops, LLMOps, Cloud platform,Model servicing optimization, Python
• Strong communication skills
• Strong programming skills
Interview Process (Is face to face required?)
• Face to face interview
Any additional information you would like to share about the project specs/nature of work:
Project Code: of Observability, Agentic AI Use cases f