... prompt evaluation. · Manage AI-assisted content creation for whitepapers, websites, and RFP ... Strong understanding of LLM frameworks (RA, or similar) and retrieval-augmented generation (RAG ...
Quick apply
... prompt evaluation. · Manage AI-assisted content creation for whitepapers, websites, and RFP ... Strong understanding of LLM frameworks (RA, or similar) and retrieval-augmented generation (RAG ...
Quick apply
... prompt evaluation. · Manage AI-assisted content creation for whitepapers, websites, and RFP ... Strong understanding of LLM frameworks (RA, or similar) and retrieval-augmented generation (RAG ...
$100K - $150K/yr
LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...
$100K - $150K/yr
LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...
Atlanta, GA · On-site
Preferred : • Experience evaluating and fine-tuning LLMs or working with RAG architectures. • ... open-source LLM, agent, or prompt engineering projects. Company : Diversity Nexus provides ...
Atlanta, GA · On-site
Preferred : • Experience evaluating and fine-tuning LLMs or working with RAG architectures. • ... open-source LLM, agent, or prompt engineering projects. Company : Diversity Nexus provides ...
Cupertino, CA · On-site
... evaluation ... A successful candidate is experienced in survey design, data annotation, LLM prompt engineering and ...
Cupertino, CA · On-site
... evaluation ... A successful candidate is experienced in survey design, data annotation, LLM prompt engineering and ...
Jersey City, NJ · On-site
LLM Interaction Designer; Conversational AI Designer; GenAI Specialist; AI Content Designer; NLP Prompt Specialist Core keywords prompt engineering, system prompt, few-shot, RAG, prompt evaluation ...
Jersey City, NJ · On-site
LLM Interaction Designer; Conversational AI Designer; GenAI Specialist; AI Content Designer; NLP Prompt Specialist Core keywords prompt engineering, system prompt, few-shot, RAG, prompt evaluation ...
San Francisco, CA · On-site
$123K - $169K/yr
... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...
San Francisco, CA · On-site
$123K - $169K/yr
... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...
San Francisco, CA · On-site +1
$123K - $169K/yr
... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...
San Francisco, CA · On-site +1
$123K - $169K/yr
... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...
OR · On-site +1
We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... the LLM prompts that power Five9 AQM's automated evaluation criteria. If you have spent years ...
OR · On-site +1
We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... the LLM prompts that power Five9 AQM's automated evaluation criteria. If you have spent years ...
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
Maintain test scripts, evaluation harnesses, simulation scenarios, and acceptancecriteria ... Strong knowledge of AI/LLM prompt engineering, including prompt structuring, testing, and ...
Maintain test scripts, evaluation harnesses, simulation scenarios, and acceptancecriteria ... Strong knowledge of AI/LLM prompt engineering, including prompt structuring, testing, and ...
We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... the LLM prompts that power Five9 AQM's automated evaluation criteria. If you have spent years ...
We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... the LLM prompts that power Five9 AQM's automated evaluation criteria. If you have spent years ...
Dallas, TX · On-site
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
Dallas, TX · On-site
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
San Ramon, CA · On-site
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
San Ramon, CA · On-site
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
Cupertino, CA · On-site
$150K - $277K/yr
... evaluation ... A successful candidate is experienced in survey design, data annotation, LLM prompt engineering and ...
Cupertino, CA · On-site
$150K - $277K/yr
... evaluation ... A successful candidate is experienced in survey design, data annotation, LLM prompt engineering and ...
Maintain test scripts, evaluation harnesses, simulation scenarios, and acceptancecriteria ... Strong knowledge of AI/LLM prompt engineering, including prompt structuring, testing, and ...
Maintain test scripts, evaluation harnesses, simulation scenarios, and acceptancecriteria ... Strong knowledge of AI/LLM prompt engineering, including prompt structuring, testing, and ...
Mclean, VA · On-site
Develop data preprocessing, feature engineering, model training, and evaluation pipelines. * Work ... Work with OpenAI/LLM APIs, prompt engineering, embeddings, vector databases, and RAG architectures
Mclean, VA · On-site
Develop data preprocessing, feature engineering, model training, and evaluation pipelines. * Work ... Work with OpenAI/LLM APIs, prompt engineering, embeddings, vector databases, and RAG architectures
Fremont, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Fremont, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Austin, TX · On-site
This role focuses on prompt optimization, RAG tuning, inference efficiency, and modern evaluation ... Required : • highly skilled LLM Optimization & Evaluation Specialist • design, optimize, and ...
Austin, TX · On-site
This role focuses on prompt optimization, RAG tuning, inference efficiency, and modern evaluation ... Required : • highly skilled LLM Optimization & Evaluation Specialist • design, optimize, and ...
Milpitas, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Milpitas, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Hayward, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Hayward, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
$7.45 - $10.53
18% of jobs
$10.53 - $13.61
2% of jobs
$14.08 is the 25th percentile. Wages below this are outliers.
$13.61 - $16.70
30% of jobs
$16.70 - $19.78
15% of jobs
$19.78 - $22.86
9% of jobs
$23.32 is the 75th percentile. Wages above this are outliers.
$22.86 - $25.94
5% of jobs
$25.94 - $29.02
3% of jobs
$29.02 - $32.10
5% of jobs
$32.10 - $35.18
4% of jobs
$35.18 - $38.26
4% of jobs
$38.26 - $41.35
3% of jobs
$7
$20
$41
An LLM Prompt Evaluation job involves assessing and optimizing prompts used in large language models (LLMs) to ensure they generate accurate, relevant, and high-quality responses. Evaluators test different prompts, analyze model outputs, and refine phrasing to improve performance. This role requires a strong understanding of AI behavior, critical thinking, and sometimes domain-specific expertise to create effective instructions for the model.
Professionals in LLM Prompt Evaluation spend their days reviewing, analyzing, and scoring the outputs of large language models based on specific prompts. This involves identifying inaccuracies, biases, or other issues in AI-generated responses, and providing detailed, constructive feedback that informs further model development. Collaboration with data scientists, machine learning engineers, and product teams is common to align evaluation efforts with organizational goals. The role also includes documenting findings, participating in regular team meetings, and staying updated on best practices in prompt design and AI ethics.
To thrive in LLM Prompt Evaluation, you need a solid understanding of natural language processing, critical thinking, and analytical skills, often supported by a background in linguistics, computer science, or related disciplines. Familiarity with annotation tools, large language model (LLM) platforms, and prompt engineering frameworks is important. Attention to detail, strong written communication, and the ability to provide objective, structured feedback are key soft skills. These qualifications ensure accurate assessments of AI-generated responses and support the continuous improvement of language models.
Cities with the most Llm Prompt Evaluation job openings:
States with the most job openings for Llm Prompt Evaluation jobs include:
The top searched job categories for Llm Prompt Evaluation jobs are:

Houston, TX • Remote
Full-time
Re-posted 26 days ago
Design, refine, and optimize system prompts, agent roles, and persona-driven prompt libraries.
· Build enterprise prompt & knowledge playbooks for reuse across consulting/workflows.
· Work with SMEs and marketing to create structured knowledge content.
· Implement Content Operations Pipelines for structured taxonomies and prompt evaluation.
· Manage AI-assisted content creation for whitepapers, websites, and RFP proposals
8+ years of experience in content strategy, marketing operations, or knowledge management, with at least 2–3 years in LLM / prompt engineering or AI content systems.
· • Strong understanding of LLM frameworks (RA, or similar) and retrieval-augmented generation (RAG).
· • Expertise in designing structured prompt frameworks, few-shot prompting, and prompt performance tuning.
· • Proficiency with content pipeline tools (Markdown, CMS, Miro, Notion, Airtable, or similar).
· • Experience building and managing enterprise knowledge systems or AI knowledge assistants.