1

Overnight Retrieval Augmented Generation Jobs (NOW HIRING)

Develop LLM-powered applications leveraging Retrieval-Augmented Generation (RAG), tool calling, and orchestration frameworks. * Build scalable APIs, microservices, and integrations supporting ...

Senior AI Technologist

Raleigh, NC · On-site

$48.75 - $63/hr

Working closely with business units, engineers, and functional teams, you will leverage applied AI technologies including large language models (LLMs), retrieval-augmented generation (RAG), AI agents ...

The ideal candidate will have a strong background in investment banking, hands-on experience with Microsoft Azure OpenAI, and expertise in Retrieval-Augmented Generation (RAG). Key Responsibilities:

Closure Technologies is seeking a AI/ML Engineer who will Implement and maintain Retrieval-Augmented Generation (RAG) pipelines and integrate Large Language Models (LLMs) into applications, supported ...

GPT, Claude • Prompt Engineering • RAG (Retrieval Augmented Generation) • AWS Cloud • Strong architectural and hands on GenAI expertise • Experience with enterprise automation and testing ...

Showing results 21-40

Overnight Retrieval Augmented Generation information

What is the difference between Overnight Retrieval Augmented Generation vs Data Scientist?

AspectOvernight Retrieval Augmented GenerationData Scientist
CredentialsTypically requires knowledge of AI, NLP, and data retrieval techniquesRequires degrees in data science, statistics, or related fields
Work EnvironmentOften in AI research labs, tech companies, or startups focusing on NLP modelsIn corporate, research, or consulting settings analyzing data and building models
Industry UsagePrimarily in AI, machine learning, and NLP industriesAcross finance, healthcare, tech, and other sectors

Overnight Retrieval Augmented Generation focuses on developing AI models that combine retrieval techniques with generative AI, often working overnight to update or improve models. Data Scientists analyze data, build predictive models, and interpret results across various industries. While both roles involve data and AI, Retrieval Augmented Generation specialists focus on model training and NLP innovations, whereas Data Scientists handle broader data analysis and modeling tasks.

More about Overnight Retrieval Augmented Generation jobs

What cities are hiring for Overnight Retrieval Augmented Generation jobs?

Cities with the most Overnight Retrieval Augmented Generation job openings:

What are the most commonly searched types of Retrieval Augmented Generation jobs?

The most popular types of Retrieval Augmented Generation jobs are:

What states have the most Overnight Retrieval Augmented Generation jobs?

States with the most job openings for Overnight Retrieval Augmented Generation jobs include:

What are popular job titles related to Overnight Retrieval Augmented Generation jobs?

For Overnight Retrieval Augmented Generation jobs, the most frequently searched job titles are:

Infographic showing various Overnight Retrieval Augmented Generation job openings in the United States as of September 2026, with employment types broken down into 2% Internship, 66% Full Time, 31% Part Time, and 1% Contract. Highlights an 63% Physical, 2% Hybrid, and 35% Remote job distribution.

Software Developer / Engineer - Philadelphia, PA (Locals Only)

Philadelphia, PA • On-site

Apetan Consulting llc
IT Services • 1 - 10 employees

$80 - $150/hr

Contractor

Re-posted 21 days ago


Job description

Software Developer / Engineer
Location: Philadelphia, PA 
Work Schedule: Hybrid 3 days on site, 2 remote


Position Overview

We are seeking a Software Developer / Engineer to help design and implement an on-premises Large Language Model (LLM) platform with Retrieval-Augmented Generation (RAG) capabilities. This role will focus on deploying open-source AI models, integrating vector databases, and building secure, enterprise-grade AI solutions in a private environment.

This is an excellent opportunity for a developer with hands-on experience in modern AI technologies who enjoys building scalable, high-performance systems.

Responsibilities

  • Deploy and optimize open-source large language models (LLMs) such as Meta Llama 3 and Mistral/Mixtral in on-premises or private environments.
  • Develop Python-based applications for LLM inference, prompt engineering, and model integration.
  • Optimize CPU-based model inference through quantization and performance tuning.
  • Design and implement Retrieval-Augmented Generation (RAG) (RAG) pipelines.
  • Configure and manage open-source vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Generate and manage embeddings while implementing metadata filtering strategies.
  • Support enterprise security requirements, including air-gapped deployments, access controls, data privacy, and audit logging.
  • Produce technical documentation, deployment guidance, and knowledge transfer materials for internal teams.
  • Build a working prototype integrating an LLM, vector database, and RAG architecture.

Required Qualifications

  • Professional experience deploying open-source LLMs (e.g., Meta Llama 3, Mistral/Mixtral) in on-premises or private environments.
  • Strong Python development experience.
  • Hands-on experience with LLM inference, prompt engineering, and AI application integration.
  • Experience optimizing CPU-based inference through model quantization and performance tuning.
  • Experience with vector databases such as Qdrant, Chroma, Milvus, or pgvector.
  • Proven experience implementing Retrieval-Augmented Generation (RAG) solutions.
  • Understanding of enterprise security, data privacy, air-gapped environments, access controls, and audit logging.

Preferred Qualifications

  • Experience with LangChain or LlamaIndex.
  • Familiarity with Docker and Kubernetes.
  • Experience with inference frameworks such as vLLM, llama.cpp, or Hugging Face Transformers.
  • Experience with Rust, Go, or C++.
  • Previous experience working in enterprise or regulated environments.

Deliverables

  • Reference architecture and deployment guidance.
  • Working prototype integrating an LLM, vector database, and RAG solution.
  • Technical documentation and knowledge transfer to internal teams.