1

Retrieval Augmented Generation Rag Jobs in Pennsylvania

Design and implement enterprise Retrieval Augmented Generation (RAG) architectures for GenAI platforms and applications. * Build and optimize semantic retrieval pipelines, vector search ...

Design and implement enterprise Retrieval Augmented Generation (RAG) architectures for GenAI platforms and applications. * Build and optimize semantic retrieval pipelines, vector search ...

Design and implement enterprise Retrieval Augmented Generation (RAG) architectures for GenAI platforms and applications. * Build and optimize semantic retrieval pipelines, vector search ...

Python AI Developer

Malvern, PA · On-site

$49.25 - $68/hr

Experience designing and implementing Retrieval-Augmented Generation (RAG) solutions * Hands-on experience with LangChain or similar AI orchestration frameworks * Experience with AWS services such as:

GenAI Ops Solution Architect

Pittsburgh, PA · On-site

$61.25 - $80.50/hr

... Retrieval Augmented Generation (RAG), Agentic AI, ModelOps, Evaluation, Observability, and AI Governance. Future duties and responsibilities Enterprise GenAI Architecture * Define and govern the ...

GenAI Ops Solution Architect

Pittsburgh, PA · On-site

$61.25 - $80.50/hr

... Retrieval Augmented Generation (RAG), Agentic AI, ModelOps, Evaluation, Observability, and AI Governance. Future duties and responsibilities Enterprise GenAI Architecture * Define and govern the ...

GenAI Ops Solution Architect

Pittsburgh, PA · On-site

$61.25 - $80.50/hr

... Retrieval Augmented Generation (RAG), Agentic AI, ModelOps, Evaluation, Observability, and AI Governance. Future duties and responsibilities Enterprise GenAI Architecture * Define and govern the ...

Computer/Data Scientist

West Mifflin, PA · On-site

$120K - $145K/yr

This position focuses on applying retrieval-augmented generation (RAG) and large language models (LLMs) to enhance information retrieval and enable the creation of detailed safety documents and ...

Computer/Data Scientist

West Mifflin, PA · On-site

$128K - $145K/yr

This position focuses on applying retrieval-augmented generation (RAG) and large language models (LLMs) to enhance information retrieval and enable the creation of detailed safety documents and ...

... Retrieval-Augmented Generation (RAG) solutions for enterprise AI use cases. • Develop and optimize prompt engineering strategies to improve AI solution performance and outcomes. • Establish and ...

next page

Showing results 1-20

Retrieval Augmented Generation Rag information

What are popular job titles related to Retrieval Augmented Generation Rag jobs in Pennsylvania?

For Retrieval Augmented Generation Rag jobs in Pennsylvania, the most frequently searched job titles are:

What job categories do people searching Retrieval Augmented Generation Rag jobs in Pennsylvania look for?

The top searched job categories for Retrieval Augmented Generation Rag jobs in Pennsylvania are:

What cities in Pennsylvania are hiring for Retrieval Augmented Generation Rag jobs?

Cities in Pennsylvania with the most Retrieval Augmented Generation Rag job openings:

Software Developer / Engineer - Philadelphia, PA (Locals Only)

Apetan Consulting llc

Philadelphia, PA • On-site

$80 - $150/hr

Contractor

Posted 29 days ago


Job description

Software Developer / Engineer
Location: Philadelphia, PA 
Work Schedule: Hybrid 3 days on site, 2 remote


Position Overview

We are seeking a Software Developer / Engineer to help design and implement an on-premises Large Language Model (LLM) platform with Retrieval-Augmented Generation (RAG) capabilities. This role will focus on deploying open-source AI models, integrating vector databases, and building secure, enterprise-grade AI solutions in a private environment.

This is an excellent opportunity for a developer with hands-on experience in modern AI technologies who enjoys building scalable, high-performance systems.

Responsibilities

  • Deploy and optimize open-source large language models (LLMs) such as Meta Llama 3 and Mistral/Mixtral in on-premises or private environments.
  • Develop Python-based applications for LLM inference, prompt engineering, and model integration.
  • Optimize CPU-based model inference through quantization and performance tuning.
  • Design and implement Retrieval-Augmented Generation (RAG) (RAG) pipelines.
  • Configure and manage open-source vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Generate and manage embeddings while implementing metadata filtering strategies.
  • Support enterprise security requirements, including air-gapped deployments, access controls, data privacy, and audit logging.
  • Produce technical documentation, deployment guidance, and knowledge transfer materials for internal teams.
  • Build a working prototype integrating an LLM, vector database, and RAG architecture.

Required Qualifications

  • Professional experience deploying open-source LLMs (e.g., Meta Llama 3, Mistral/Mixtral) in on-premises or private environments.
  • Strong Python development experience.
  • Hands-on experience with LLM inference, prompt engineering, and AI application integration.
  • Experience optimizing CPU-based inference through model quantization and performance tuning.
  • Experience with vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Proven experience implementing Retrieval-Augmented Generation (RAG) solutions.
  • Understanding of enterprise security, data privacy, air-gapped environments, access controls, and audit logging.

Preferred Qualifications

  • Experience with LangChain or LlamaIndex.
  • Familiarity with Docker and Kubernetes.
  • Experience with inference frameworks such as vLLM, llama.cpp, or Hugging Face Transformers.
  • Experience with Rust, Go, or C++.
  • Previous experience working in enterprise or regulated environments.

Deliverables

  • Reference architecture and deployment guidance.
  • Working prototype integrating an LLM, vector database, and RAG solution.
  • Technical documentation and knowledge transfer to internal teams.