... prompt evaluation. · Manage AI-assisted content creation for whitepapers, websites, and RFP ... Strong understanding of LLM frameworks (RA, or similar) and retrieval-augmented generation (RAG ...
Quick apply
... prompt evaluation. · Manage AI-assisted content creation for whitepapers, websites, and RFP ... Strong understanding of LLM frameworks (RA, or similar) and retrieval-augmented generation (RAG ...
Quick apply
... prompt evaluation. · Manage AI-assisted content creation for whitepapers, websites, and RFP ... Strong understanding of LLM frameworks (RA, or similar) and retrieval-augmented generation (RAG ...
Ashburn, VA · On-site
$100K - $150K/yr
LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...
Ashburn, VA · On-site
$100K - $150K/yr
LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...
$100K - $150K/yr
LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...
$100K - $150K/yr
LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...
Sterling, VA · On-site
$100K - $150K/yr
LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...
New
Sterling, VA · On-site
$100K - $150K/yr
LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...
New
Cupertino, CA · On-site
... evaluation ... A successful candidate is experienced in survey design, data annotation, LLM prompt engineering and ...
Cupertino, CA · On-site
... evaluation ... A successful candidate is experienced in survey design, data annotation, LLM prompt engineering and ...
Jersey City, NJ · On-site
LLM Interaction Designer; Conversational AI Designer; GenAI Specialist; AI Content Designer; NLP Prompt Specialist Core keywords prompt engineering, system prompt, few-shot, RAG, prompt evaluation ...
Jersey City, NJ · On-site
LLM Interaction Designer; Conversational AI Designer; GenAI Specialist; AI Content Designer; NLP Prompt Specialist Core keywords prompt engineering, system prompt, few-shot, RAG, prompt evaluation ...
San Francisco, CA · On-site
$123K - $169K/yr
... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...
San Francisco, CA · On-site
$123K - $169K/yr
... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...
San Francisco, CA · On-site +1
$123K - $169K/yr
... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...
San Francisco, CA · On-site +1
$123K - $169K/yr
... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...
We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... the LLM prompts that power Five9 AQM's automated evaluation criteria. If you have spent years ...
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
Quick apply
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
Quick apply
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
Quick apply
Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...
We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... the LLM prompts that power Five9 AQM's automated evaluation criteria. If you have spent years ...
We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... the LLM prompts that power Five9 AQM's automated evaluation criteria. If you have spent years ...
... evaluation ... A successful candidate is experienced in survey design, data annotation, LLM prompt engineering and ...
... evaluation ... A successful candidate is experienced in survey design, data annotation, LLM prompt engineering and ...
Maintain test scripts, evaluation harnesses, simulation scenarios, and acceptancecriteria ... Strong knowledge of AI/LLM prompt engineering, including prompt structuring, testing, and ...
Maintain test scripts, evaluation harnesses, simulation scenarios, and acceptancecriteria ... Strong knowledge of AI/LLM prompt engineering, including prompt structuring, testing, and ...
Maintain test scripts, evaluation harnesses, simulation scenarios, and acceptancecriteria ... Strong knowledge of AI/LLM prompt engineering, including prompt structuring, testing, and ...
Maintain test scripts, evaluation harnesses, simulation scenarios, and acceptancecriteria ... Strong knowledge of AI/LLM prompt engineering, including prompt structuring, testing, and ...
Fremont, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Fremont, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Milpitas, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Milpitas, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Mountain View, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Mountain View, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Cupertino, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
Cupertino, CA · On-site
$143K - $286K/yr
Define the strategy for AI-native experimentation and evaluation, including LLM eval frameworks, prompt evaluation, golden datasets, rubric design, human-in-the-loop review, LLM-as-a-judge ...
$7.45 - $10.53
18% of jobs
$10.53 - $13.61
2% of jobs
$14.08 is the 25th percentile. Wages below this are outliers.
$13.61 - $16.70
30% of jobs
$16.70 - $19.78
15% of jobs
$19.78 - $22.86
9% of jobs
$23.32 is the 75th percentile. Wages above this are outliers.
$22.86 - $25.94
5% of jobs
$25.94 - $29.02
3% of jobs
$29.02 - $32.10
5% of jobs
$32.10 - $35.18
4% of jobs
$35.18 - $38.26
4% of jobs
$38.26 - $41.35
3% of jobs
$7
$20
$41
To thrive in LLM Prompt Evaluation, you need a solid understanding of natural language processing, critical thinking, and analytical skills, often supported by a background in linguistics, computer science, or related disciplines. Familiarity with annotation tools, large language model (LLM) platforms, and prompt engineering frameworks is important. Attention to detail, strong written communication, and the ability to provide objective, structured feedback are key soft skills. These qualifications ensure accurate assessments of AI-generated responses and support the continuous improvement of language models.
An LLM Prompt Evaluation job involves assessing and optimizing prompts used in large language models (LLMs) to ensure they generate accurate, relevant, and high-quality responses. Evaluators test different prompts, analyze model outputs, and refine phrasing to improve performance. This role requires a strong understanding of AI behavior, critical thinking, and sometimes domain-specific expertise to create effective instructions for the model.
Professionals in LLM Prompt Evaluation spend their days reviewing, analyzing, and scoring the outputs of large language models based on specific prompts. This involves identifying inaccuracies, biases, or other issues in AI-generated responses, and providing detailed, constructive feedback that informs further model development. Collaboration with data scientists, machine learning engineers, and product teams is common to align evaluation efforts with organizational goals. The role also includes documenting findings, participating in regular team meetings, and staying updated on best practices in prompt design and AI ethics.

Full-time
Re-posted 5 days ago
Design, refine, and optimize system prompts, agent roles, and persona-driven prompt libraries.
· Build enterprise prompt & knowledge playbooks for reuse across consulting/workflows.
· Work with SMEs and marketing to create structured knowledge content.
· Implement Content Operations Pipelines for structured taxonomies and prompt evaluation.
· Manage AI-assisted content creation for whitepapers, websites, and RFP proposals
8+ years of experience in content strategy, marketing operations, or knowledge management, with at least 2–3 years in LLM / prompt engineering or AI content systems.
· • Strong understanding of LLM frameworks (RA, or similar) and retrieval-augmented generation (RAG).
· • Expertise in designing structured prompt frameworks, few-shot prompting, and prompt performance tuning.
· • Proficiency with content pipeline tools (Markdown, CMS, Miro, Notion, Airtable, or similar).
· • Experience building and managing enterprise knowledge systems or AI knowledge assistants.