1

Llm Prompt Evaluation Jobs (NOW HIRING)

LLM Prompt Specialist

Ashburn, VA · On-site

$100K - $150K/yr

LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...

LLM Prompt Specialist

$100K - $150K/yr

LLM Prompt Specialist - Remote Bright Vision Technologies is a technology consulting and software ... Experience designing evaluation pipelines for non-deterministic systems. * Strong Python skills and ...

LLM Interaction Designer; Conversational AI Designer; GenAI Specialist; AI Content Designer; NLP Prompt Specialist Core keywords prompt engineering, system prompt, few-shot, RAG, prompt evaluation ...

Senior Prompt Engineering Engineer

San Francisco, CA · On-site

$123K - $169K/yr

... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...

Senior Prompt Engineering Engineer

San Francisco, CA · On-site +1

$123K - $169K/yr

... on LLM prompt or agent engineering. - Deep familiarity with Claude, GPT, Gemini, or open-source families, as well as evaluation stacks such as Ragas, DeepEval, LangSmith, Weights & Biases, etc ...

We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... the LLM prompts that power Five9 AQM's automated evaluation criteria. If you have spent years ...

Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...

Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...

Design, author, and iteratively optimize LLM evaluation prompts within Five9 AQM, translating quality management frameworks into accurate, testable prompt logic. * Partner with customers and internal ...

We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... the LLM prompts that power Five9 AQM's automated evaluation criteria. If you have spent years ...

next page

Showing results 1-20

Llm Prompt Evaluation information

See salary details

$7

$20

$41

How much do llm prompt evaluation jobs pay per hour?

As of Aug 8, 2026, the average hourly pay for llm prompt evaluation in the United States is $20.89, according to ZipRecruiter salary data. Most workers in this role earn between $14.42 and $27.64 per hour, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive in the LLM prompt evaluation position?

To thrive in LLM Prompt Evaluation, you need a solid understanding of natural language processing, critical thinking, and analytical skills, often supported by a background in linguistics, computer science, or related disciplines. Familiarity with annotation tools, large language model (LLM) platforms, and prompt engineering frameworks is important. Attention to detail, strong written communication, and the ability to provide objective, structured feedback are key soft skills. These qualifications ensure accurate assessments of AI-generated responses and support the continuous improvement of language models.

What is an LLM prompt evaluation?

An LLM Prompt Evaluation job involves assessing and optimizing prompts used in large language models (LLMs) to ensure they generate accurate, relevant, and high-quality responses. Evaluators test different prompts, analyze model outputs, and refine phrasing to improve performance. This role requires a strong understanding of AI behavior, critical thinking, and sometimes domain-specific expertise to create effective instructions for the model.

What does an LLM prompt evaluation do?

Professionals in LLM Prompt Evaluation spend their days reviewing, analyzing, and scoring the outputs of large language models based on specific prompts. This involves identifying inaccuracies, biases, or other issues in AI-generated responses, and providing detailed, constructive feedback that informs further model development. Collaboration with data scientists, machine learning engineers, and product teams is common to align evaluation efforts with organizational goals. The role also includes documenting findings, participating in regular team meetings, and staying updated on best practices in prompt design and AI ethics.

More about Llm Prompt Evaluation jobs
What cities are hiring for Llm Prompt Evaluation jobs? Cities with the most Llm Prompt Evaluation job openings:
What states have the most Llm Prompt Evaluation jobs? States with the most job openings for Llm Prompt Evaluation jobs include:
Infographic showing various Llm Prompt Evaluation job openings in the United States as of August 2026, with employment types broken down into 2% As Needed, 79% Full Time, 17% Part Time, and 2% Contract. Highlights an 90% Physical, 2% Hybrid, and 8% Remote job distribution, with an average salary of $43,449 per year, or $20.9 per hour.

Senior Prompt & Content Engineering Lead

Bay Area TeK Solutions LLC

Houston, TX • Remote

Full-time

Re-posted 5 days ago


Job description

Design, refine, and optimize system prompts, agent roles, and persona-driven prompt libraries.

· Build enterprise prompt & knowledge playbooks for reuse across consulting/workflows.

· Work with SMEs and marketing to create structured knowledge content.

· Implement Content Operations Pipelines for structured taxonomies and prompt evaluation.

· Manage AI-assisted content creation for whitepapers, websites, and RFP proposals

8+ years of experience in content strategy, marketing operations, or knowledge management, with at least 2–3 years in LLM / prompt engineering or AI content systems.

· • Strong understanding of LLM frameworks (RA, or similar) and retrieval-augmented generation (RAG).

· • Expertise in designing structured prompt frameworks, few-shot prompting, and prompt performance tuning.

· • Proficiency with content pipeline tools (Markdown, CMS, Miro, Notion, Airtable, or similar).

· • Experience building and managing enterprise knowledge systems or AI knowledge assistants.