1

Llm Prompt Review Jobs (NOW HIRING)

Senior WEM AQM Prompt Engineer

Dallas, TX

$103K - $142K/yr

... reviews and providing structured feedback. * Drive continuous improvement of the team's prompt design process based on performance data, customer outcomes, and LLM platform developments. Key ...

Establish review, testing, and gating processes for AI-generated code across frontend, backend, IaC ... Extend DLP coverage to LLM prompt submissions and completions. * Detection, monitoring, and ...

... LLM prompt engineering. * Deep hands-on experience with Claude, GPT-4, or similar large language ... For further information, please review the Know Your Rights notice from the Department of Labor.

LLM Engineer

Houston, TX · On-site

$120K - $130K/yr

This role focuses on using artificial intelligence tools, prompt engineering techniques, and ... Review, validate, and refine AI-generated content to ensure accuracy and data integrity * Support ...

Design and maintain a safety evaluation framework-adversarial prompt sets, scenario-based test ... Conduct safety reviews of reinforcement-learning (RL) environments and trajectory data, partnering ...

Design and maintain a safety evaluation framework--adversarial prompt sets, scenario-based test ... Conduct safety reviews of reinforcement-learning (RL) environments and trajectory data, partnering ...

LLM Prompt Optimization - Design, test, and refine prompts to get high-quality outputs from large ... Deliver actionable insights, performance reviews, and monthly deliverables (e.g., reports, strategy ...

New

LLM Prompt Optimization - Design, test, and refine prompts to get high-quality outputs from large ... Deliver actionable insights, performance reviews, and monthly deliverables (e.g., reports, strategy ...

New

Mentor engineers through design reviews, code reviews, and technical guidance * Quickly evaluate ... LLM Prompt Engineering & Fine-Tuning: designing and evaluating effective prompts, system ...

LLM Prompt Optimization - Design, test, and refine prompts to get high-quality outputs from large ... Deliver actionable insights, performance reviews, and monthly deliverables (e.g., reports, strategy ...

... virtual assistants using LLM and generative AI technologies. * Software Engineering ... Conduct quality assurance reviews to validate AI interaction quality and system performance

LLM Prompt Optimization - Design, test, and refine prompts to get high-quality outputs from large ... Deliver actionable insights, performance reviews, and monthly deliverables (e.g., reports, strategy ...

Create and maintain engineering documentation required for customer security reviews, compliance ... LLM prompt is just Schrödinger's cat waiting to be observed, and knows too many random facts about ...

... and code review. * Ensure solutions align with enterprise standards for security, privacy ... Experience with LLM prompt engineering and architectural patterns for LLM systems (e.g., retrieval ...

... virtual assistants using LLM and generative AI technologies. * Software Engineering ... reviews, and prompt optimization to improve conversational accuracy, usability, and overall AI ...

Showing results 21-40

Llm Prompt Review information

See salary details

$37

$60

$77

How much do llm prompt review jobs pay per hour?

As of Aug 13, 2026, the average hourly pay for llm prompt review in the United States is $60.90, according to ZipRecruiter salary data. Most workers in this role earn between $54.09 and $69.71 per hour, depending on experience, location, and employer.

What are some common challenges faced by professionals in LLM prompt review roles, and how can they be managed?

Professionals in LLM Prompt Review roles often encounter challenges such as ensuring prompt clarity, mitigating bias, and maintaining consistency across large volumes of prompts. Balancing creativity with precision is essential, as even small changes can significantly impact model outputs. To manage these challenges, reviewers typically rely on established guidelines, peer collaboration, regular calibration sessions, and continuous feedback from model performance metrics. Staying updated on best practices and working closely with data scientists and prompt engineers also helps maintain high-quality outputs.

What is an LLM prompt reviewer?

An LLM Prompt Reviewer is a professional responsible for evaluating, refining, and optimizing prompts used with large language models (LLMs) like GPT-4. Their main goal is to ensure that prompts elicit accurate, useful, and safe responses from the AI. This role involves understanding both the technical and linguistic aspects of prompts, testing various phrasings, and documenting best practices. LLM Prompt Reviewers often collaborate with data scientists, AI trainers, and product teams to improve prompt quality and user experience.

What is the difference between Llm Prompt Review vs Data Annotator?

AspectLlm Prompt ReviewData Annotator
CredentialsBasic understanding of AI and NLP conceptsTypically high school diploma or equivalent, sometimes specialized training
Work EnvironmentRemote or office-based, focused on AI projectsRemote or on-site, working with datasets and labeling tools
Industry UsageUsed in AI development, NLP, and machine learning projectsUsed across various industries for data preparation and labeling
Search & Comparison IntentUnderstanding roles related to AI prompt evaluationComparing data labeling and annotation roles

While both roles involve working with data and AI, Llm Prompt Review focuses on evaluating and refining AI prompts, whereas Data Annotator involves labeling data for machine learning models. The roles differ mainly in their specific tasks and required skills, but both are essential in AI development workflows.

What are the key skills and qualifications needed to thrive as an LLM prompt reviewer, and why are they important?

To thrive as an LLM Prompt Reviewer, you need a strong background in linguistics, critical thinking, and AI language model behavior, often supported by experience in content moderation or NLP. Familiarity with prompt engineering tools, annotation platforms, and basic understanding of large language model systems is typically required. Attention to detail, analytical skills, and clear written communication make someone stand out in this position. These skills ensure the creation and evaluation of high-quality prompts that drive accurate, safe, and useful AI model outputs.
More about Llm Prompt Review jobs
What cities are hiring for Llm Prompt Review jobs? Cities with the most Llm Prompt Review job openings:
What states have the most Llm Prompt Review jobs? States with the most job openings for Llm Prompt Review jobs include:
Infographic showing various Llm Prompt Review job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 83% Full Time, 13% Part Time, and 3% Contract. Highlights an 90% Physical, 3% Hybrid, and 7% Remote job distribution, with an average salary of $126,666 per year, or $60.9 per hour.

Senior WEM AQM Prompt Engineer

Five9

Dallas, TX

$103K - $142K/yr

Full-time

Posted 22 days ago


Job description

Join us in bringing joy to customer experience. Five9 is a leading provider of cloud contact center software, bringing the power of cloud innovation to customers worldwide.

Living our values everyday results in our team-first culture and enables us to innovate, grow, and thrive while enjoying the journey together. We celebrate diversity and foster an inclusive environment, empowering our employees to be their authentic selves.


Five9 is one of the world's leading cloud contact center platforms. Our AI-powered Automated Quality Management (AQM) product is transforming how enterprises measure, evaluate, and improve agent and interaction quality at scale. As we build out our dedicated AQM Prompt Engineering capability within Specialized Services, we are hiring a Senior WEM AQM Prompt Engineer to lead and shape this practice.


This senior role goes beyond hands-on prompt design. You will be the technical authority for Five9 AQM prompt engineering — setting standards, building reusable assets, mentoring the Prompt Engineer team, and acting as the go-to expert for complex customer escalations and pre-sales engagements. If you have spent a significant portion of your career building sophisticated Speech Analytics programs in Verint, NICE, or Genesys — and you have developed serious prompt engineering skills alongside that domain expertise — this role was designed for you.


Why Your Speech Analytics Background is the Foundation for This Role
Building effective AQM prompts follows the same disciplined, iterative methodology that defines expert-level Speech Analytics category design:

Speech Analytics / Verint AQMFive9 AQM Prompt EngineeringDefine category intent & scopeDefine prompt intent & evaluation scopeSelect & tag representative call samplesSelect evaluation call samplesIteratively tune keyword / phrase setsIteratively refine prompt wording & logicTest against live call trafficTest prompt outputs against real callsValidate recall & precision metricsValidate pass / fail accuracy metricsDocument & hand off to QA teamDocument prompts for QA workflow

At the senior level, this role adds a leadership and practice-building dimension: you will set the standards the rest of the Prompt Engineering team follows and serve as the technical authority across the most complex customer engagements.


Key Responsibilities — Practice Leadership
• Serve as the technical lead and subject matter authority for Five9 AQM Prompt Engineering within Specialized Services.
• Define and own the prompt engineering methodology, standards, quality gates, and best practice documentation for the team.
• Build and maintain a curated, reusable prompt library of validated evaluation templates spanning key verticals and QM use cases.
• Mentor and technically guide the WEM AQM Prompt Engineer team (x3), conducting prompt reviews and providing structured feedback.
• Drive continuous improvement of the team's prompt design process based on performance data, customer outcomes, and LLM platform developments.


Key Responsibilities — Customer Delivery
• Lead prompt engineering on the most complex and high-value AQM customer engagements, from initial scoping through production deployment.
• Act as the senior escalation point for prompt performance issues, failure analysis, and remediation on live customer deployments.
• Partner with customers at senior stakeholder level to translate QM strategy and compliance requirements into precise, measurable AQM evaluation criteria.
• Design and execute prompt validation frameworks — including representative call sampling, precision/recall analysis, and A/B prompt testing — to ensure accuracy before production release.


Key Responsibilities — Product & Pre-Sales Collaboration
• Serve as the Prompt Engineering SME in pre-sales engagements, demonstrating Five9 AQM capabilities and advising prospective customers on prompt-driven evaluation design.
• Provide structured product feedback to the Five9 AQM product team based on field experience, identifying gaps, limitations, and enhancement opportunities.
• Develop and deliver internal training and enablement for consultants and implementation staff on AQM prompt engineering techniques.
• Represent the Prompt Engineering practice in cross-functional forums and contribute to go-to-market materials, customer case studies, and thought leadership content.


Key Requirements
• Minimum 5 years' hands-on experience as a Speech Analytics Lead, Senior QM Analyst, or WEM Consultant with deep expertise in category/topic design, program governance, and optimization in Verint (Categories/Category Sets), NICE CXne (Category Sets), Genesys (Topics), or equivalent — at a seniority level sufficient to have owned or led an analytics program.
• Formal prompt engineering qualification or a verifiable, substantive portfolio of production-grade prompt design work, including evidence of iterative refinement, performance measurement, and documentation.
• Demonstrable understanding of LLM behavior — including how prompt structure, instruction framing, persona assignment, chain-of-thought, and few-shot examples affect model outputs.
• Strong background in Quality Management design: experience building evaluation frameworks, scorecard design, calibration processes, and QM governance.
• Experience leading or mentoring technical team members in a delivery or consulting environment.
• Proven track record of translating complex business requirements into structured, testable, and scalable technical solutions.
• Excellent written and verbal communication skills; comfortable presenting to senior client stakeholders and internal leadership.
• BA/BS or equivalent experience; advanced degree in a relevant discipline is advantageous.


Preferred Qualifications
• Formal prompt engineering certification: DeepLearning.AI Prompt Engineering for Developers, Anthropic Prompt Engineering, OpenAI Prompt Engineering, DAIR.AI Prompt Engineering Guide, or equivalent industry-recognized program.
• Direct experience with one or more leading LLM platforms at an advanced level (OpenAI GPT-4+, Anthropic Claude, Google Gemini, Meta Llama, or similar), including API usage, system prompt design, and evaluation harness development.
• Prior hands-on experience with Five9 AQM or another AI-native QM solution (e.g., NICE Enlighten QM, Genesys Cloud AI Quality, Medallia, Qualtrics).
• Experience contributing to or owning a Speech Analytics or QM center of excellence — including methodology documentation, internal training, and cross-client best practice standardization.
• Familiarity with prompt evaluation frameworks (e.g., LLM-as-judge, human-in-the-loop evaluation pipelines, or equivalent).
• Exposure to NLP concepts — tokenization, embeddings, semantic similarity — sufficient to reason about why a prompt succeeds or fails.
• Experience with JSON, Python scripting, or no-code automation tools as applied to prompt testing and workflow automation.
• Background in contact center compliance-driven QM (financial services, healthcare, utilities) where regulatory accuracy requirements add additional complexity to evaluation design.


Key Competencies
• Technical Authority — recognized internally and externally as the go-to expert on LLM prompt design for contact Centre quality evaluation.
• Analytical Rigor — applies statistical thinking and structured measurement to validate and continuously improve prompt performance.
• Linguistic Mastery — exceptional command of written instruction; understands how nuance, framing, and ambiguity in prompt language directly influence AI outputs.
• Leadership & Coaching — ability to elevate the capability of others through structured mentoring, peer review, and knowledge transfer.
• Strategic Customer Engagement — trusted advisor to senior client stakeholders on the use of AI in quality management.
• Innovation Orientation — proactively monitors the LLM landscape and identifies opportunities to improve Five9 AQM prompt design practices.
• Operational Excellence — drives the team toward repeatable, scalable processes without sacrificing quality or creativity.


About Five9 WEM Specialized Services
The Specialized Services team within Five9 Professional Services delivers expert-led implementation, optimization, and advisory engagements across the Five9 WEM portfolio. Our AQM Prompt Engineering practice is a newly established Centre of excellence, designed to ensure that customers realize the full potential of Five9's AI-powered quality management capabilities. As the founding senior hire into this practice, the Senior WEM AQM Prompt Engineer will have a rare opportunity to define how prompt engineering is done at Five9 — and to build something genuinely new.

Work Location: This role is fully remote for candidates who reside outside the 30 mile radius of one of our offices. For candidates who reside within a 30 mile radius of one of our offices, this role is Hybrid and would require 3 days a week (T, W, TH) in office.


As part of our continued commitment to diversity, equity, and inclusion, Five9 supports pay transparency during the entire recruitment process. Actual compensation packages are based on several factors that are unique to each candidate including, but not limited to: skill set, depth of experience, certifications, and specific work location. The range displayed reflects the minimum and maximum target for new hire salaries for the job across the United States. Your recruiter can share more about the specific compensation package during your hiring process.

Additionally, the total compensation package for this position may also include an annual performance bonus, stock, and/or other applicable incentive compensation plans.

Our total reward package also includes:

  • Health, dental, and vision coverage, beginning on the first day of employment. Five9 covers 100% of the employee portion of the health, dental and vision coverage and shares a high portion of the dependent cost. We also offer Short & Long-Term Disability, Basic Life Insurance, and a 401k saving plan with employer matching.
  • Access to an innovative mental health support platform that offers personalized care and resources in areas such as: therapy, coaching and self-guided mindfulness exercises for all covered employees and their covered dependents.
  • Generous employee stock purchase plan.
  • Paid Time Off, Company paid holidays, paid volunteer hours and 12 weeks paid parental leave.

All compensation and benefits are subject to the requirements and restrictions set forth in the applicable plan documents and any written agreements between the parties.

The US base salary range for this role is below.
$88,900—$247,700 USD

Five9 embraces diversity and is committed to building a team that represents a variety of backgrounds, perspectives, and skills.  The more inclusive we are, the better we are.  Five9 is an equal opportunity employer.


View our privacy policy, including our privacy notice to California residents here: https://www.five9.com/pt-pt/legal.

Note: Five9 will never request that an applicant send money as a prerequisite for commencing employment with Five9.