1

Retrieval Augmented Generation Jobs in Philadelphia, PA

Python AI Developer

Malvern, PA · On-site

$49.25 - $68/hr

Experience designing and implementing Retrieval-Augmented Generation (RAG) solutions * Hands-on experience with LangChain or similar AI orchestration frameworks * Experience with AWS services such as:

Hands-on experience designing and implementing retrieval-augmented generation (RAG) workflows * Understanding of: * LLM integration and orchestration * Embeddings and vector search * Retrieval ...

Senior Python AI Engineer

Mount Laurel, NJ · On-site

$120K - $161K/yr

Have prior knowledge & hands on Experience with - Building production-ready AI services and APIs using Python LLM integration Prompt engineering frameworks RAG (Retrieval-Augmented Generation ...

Hands-on experience designing and implementing retrieval-augmented generation (RAG) workflows * Understanding of: * LLM integration and orchestration * Embeddings and vector search * Retrieval ...

Software Engineer

Newtown, PA · On-site +1

$108K - $115K/yr

Help build and integrate Retrieval-Augmented Generation (RAG) pipelines with vector databases (e.g., Azure AI Search, Pinecone, pgvector) to ground AI outputs in enterprise data. * Apply Responsible ...

Hands-on experience designing and implementing retrieval-augmented generation (RAG) workflows * Understanding of: * LLM integration and orchestration * Embeddings and vector search * Retrieval ...

Software Engineer

Newtown, PA · On-site +1

$108K - $115K/yr

Help build and integrate Retrieval-Augmented Generation (RAG) pipelines with vector databases (e.g., Azure AI Search, Pinecone, pgvector) to ground AI outputs in enterprise data. * Apply Responsible ...

Mistral). • Implement RAG (Retrieval-Augmented Generation); vector databases; embeddings; and prompt‐engineering strategies. • Collaborate with product; engineering; and domain teams to ...

Hands-on experience designing or implementing Retrieval-Augmented Generation (RAG) solutions. * Experience with AI standards and frameworks such as Model Context Protocol (MCP), Agent2Agent (A2A ...

next page

Showing results 1-20

Retrieval Augmented Generation information

What is a retrieval augmented generation?

A Retrieval Augmented Generation (RAG) job typically involves developing and optimizing AI systems that enhance text generation by incorporating external knowledge retrieved from relevant sources. Professionals in this field work on integrating retrieval mechanisms with large language models to improve the relevance, accuracy, and factual grounding of generated content. Common responsibilities include designing retrieval systems, fine-tuning language models, optimizing performance, and ensuring the seamless integration of factual data into AI-generated text. This role is highly interdisciplinary, involving expertise in natural language processing (NLP), machine learning, and information retrieval.

What does a retrieval augmented generation engineer do?

A Retrieval Augmented Generation engineer typically spends their day designing and implementing systems that combine information retrieval with advanced generative models, such as large language models. This includes fine-tuning models, integrating external data sources, developing vector search pipelines, and evaluating output quality. Collaboration with data scientists, machine learning engineers, and product teams is common to ensure the solutions meet user requirements and scale effectively. Additionally, RAG engineers often troubleshoot issues, monitor model performance in production, and stay informed about the latest advancements in AI and information retrieval.

What skills and qualifications are needed for retrieval augmented generation?

To thrive in a Retrieval Augmented Generation (RAG) engineering role, you need a solid background in machine learning, natural language processing (NLP), and experience with scalable information retrieval systems, typically supported by a relevant degree in computer science or a related field. Familiarity with tools such as Python, PyTorch or TensorFlow, vector databases, and search platforms like Elasticsearch is essential, along with practical experience deploying and tuning RAG pipelines. Strong problem-solving skills, a collaborative mindset, and effective communication abilities set outstanding professionals apart in this field. These competencies are crucial for designing, implementing, and optimizing hybrid retrieval-generation AI systems that address complex, real-world information needs.

What are the most commonly searched types of Retrieval Augmented Generation jobs in Philadelphia, PA?

The most popular types of Retrieval Augmented Generation jobs in Philadelphia, PA are:

What are popular job titles related to Retrieval Augmented Generation jobs in Philadelphia, PA?

For Retrieval Augmented Generation jobs in Philadelphia, PA, the most frequently searched job titles are:

What job categories do people searching Retrieval Augmented Generation jobs in Philadelphia, PA look for?

The top searched job categories for Retrieval Augmented Generation jobs in Philadelphia, PA are:

What cities near Philadelphia, PA are hiring for Retrieval Augmented Generation jobs?

Cities near Philadelphia, PA with the most Retrieval Augmented Generation job openings:

Infographic showing various Retrieval Augmented Generation job openings in Philadelphia, PA as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% Remote job distribution.

Software Developer / Engineer - Philadelphia, PA (Locals Only)

Apetan Consulting llc

Philadelphia, PA • On-site

$80 - $150/hr

Contractor

Posted 28 days ago


Job description

Software Developer / Engineer
Location: Philadelphia, PA 
Work Schedule: Hybrid 3 days on site, 2 remote


Position Overview

We are seeking a Software Developer / Engineer to help design and implement an on-premises Large Language Model (LLM) platform with Retrieval-Augmented Generation (RAG) capabilities. This role will focus on deploying open-source AI models, integrating vector databases, and building secure, enterprise-grade AI solutions in a private environment.

This is an excellent opportunity for a developer with hands-on experience in modern AI technologies who enjoys building scalable, high-performance systems.

Responsibilities

  • Deploy and optimize open-source large language models (LLMs) such as Meta Llama 3 and Mistral/Mixtral in on-premises or private environments.
  • Develop Python-based applications for LLM inference, prompt engineering, and model integration.
  • Optimize CPU-based model inference through quantization and performance tuning.
  • Design and implement Retrieval-Augmented Generation (RAG) (RAG) pipelines.
  • Configure and manage open-source vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Generate and manage embeddings while implementing metadata filtering strategies.
  • Support enterprise security requirements, including air-gapped deployments, access controls, data privacy, and audit logging.
  • Produce technical documentation, deployment guidance, and knowledge transfer materials for internal teams.
  • Build a working prototype integrating an LLM, vector database, and RAG architecture.

Required Qualifications

  • Professional experience deploying open-source LLMs (e.g., Meta Llama 3, Mistral/Mixtral) in on-premises or private environments.
  • Strong Python development experience.
  • Hands-on experience with LLM inference, prompt engineering, and AI application integration.
  • Experience optimizing CPU-based inference through model quantization and performance tuning.
  • Experience with vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Proven experience implementing Retrieval-Augmented Generation (RAG) solutions.
  • Understanding of enterprise security, data privacy, air-gapped environments, access controls, and audit logging.

Preferred Qualifications

  • Experience with LangChain or LlamaIndex.
  • Familiarity with Docker and Kubernetes.
  • Experience with inference frameworks such as vLLM, llama.cpp, or Hugging Face Transformers.
  • Experience with Rust, Go, or C++.
  • Previous experience working in enterprise or regulated environments.

Deliverables

  • Reference architecture and deployment guidance.
  • Working prototype integrating an LLM, vector database, and RAG solution.
  • Technical documentation and knowledge transfer to internal teams.