Oversee structured A/B testing initiatives and data analysis to iteratively optimize agent ... AI, Large Language Models (LLMs), and deep prompt engineering strategies within an enterprise AWS ...
Quick apply
Oversee structured A/B testing initiatives and data analysis to iteratively optimize agent ... AI, Large Language Models (LLMs), and deep prompt engineering strategies within an enterprise AWS ...
Quick apply
Oversee structured A/B testing initiatives and data analysis to iteratively optimize agent ... AI, Large Language Models (LLMs), and deep prompt engineering strategies within an enterprise AWS ...
$69K - $89K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
$69K - $89K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
$70K - $90K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
$70K - $90K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
This role is responsible for defining enterprise prompt strategy, AI personas, and multi-agent ... Lead A/B testing initiatives for prompts, conversational flows, personas, and response strategies.
This role is responsible for defining enterprise prompt strategy, AI personas, and multi-agent ... Lead A/B testing initiatives for prompts, conversational flows, personas, and response strategies.
$69K - $89K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
$69K - $89K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
$67K - $86K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
$67K - $86K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
$70K - $90K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
$70K - $90K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
Tysons Corner, VA · On-site
$72K - $93K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
Tysons Corner, VA · On-site
$72K - $93K/yr
The Prompt + Skills Engineer is the hands-on builder in Cherry Bekaert's AI Center of Excellence ... Conducts quality assurance testing on AI outputs, including accuracy validation, edge case testing ...
Austin, TX · On-site +1
... testing. * Share insights and learnings with the broader product and AI teams. Requirements ... Strong skills in prompt engineering techniques (e.g., few-shot, chain-of-thought, RAG pipelines)
Austin, TX · On-site +1
... testing. * Share insights and learnings with the broader product and AI teams. Requirements ... Strong skills in prompt engineering techniques (e.g., few-shot, chain-of-thought, RAG pipelines)
Austin, TX · On-site
... testing. * Share insights and learnings with the broader product and AI teams. Requirements ... Strong skills in prompt engineering techniques (e.g., few-shot, chain-of-thought, RAG pipelines)
Quick apply
Austin, TX · On-site
... testing. * Share insights and learnings with the broader product and AI teams. Requirements ... Strong skills in prompt engineering techniques (e.g., few-shot, chain-of-thought, RAG pipelines)
Prompt design & iteration for generative AI systems -- 2+ years (or equivalent hands-on experience ... Analytical problem solving (testing outputs, identifying failure patterns, improving systems) -- 2+ ...
Quick apply
Prompt design & iteration for generative AI systems -- 2+ years (or equivalent hands-on experience ... Analytical problem solving (testing outputs, identifying failure patterns, improving systems) -- 2+ ...
OR · On-site +1
We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... and applied AI - designing, testing, and optimizing the LLM prompts that power Five9 AQM ...
OR · On-site +1
We are seeking three WEM AQM Prompt Engineers to join this new capability team. These roles sit at ... and applied AI - designing, testing, and optimizing the LLM prompts that power Five9 AQM ...
Experience with LLM observability and eval tooling - Confident AI, LangFuse, LangSmith, or similar ... Experience with red teaming or adversarial prompt testing * Understanding of prompt injection risks ...
Experience with LLM observability and eval tooling - Confident AI, LangFuse, LangSmith, or similar ... Experience with red teaming or adversarial prompt testing * Understanding of prompt injection risks ...
Denver, MO · Remote
Test RAG pipelines, prompt engineering, and AI agent orchestration. Execute API testing using Postman, Swagger, and REST APIs. Develop and maintain automation scripts using Python, Selenium ...
Denver, MO · Remote
Test RAG pipelines, prompt engineering, and AI agent orchestration. Execute API testing using Postman, Swagger, and REST APIs. Develop and maintain automation scripts using Python, Selenium ...
Denver, MO · On-site +1
Prompt Engineering Validation * AI Model Validation & Evaluation * API Testing (Postman, REST APIs, Swagger) * Python * Test Automation (Selenium / Playwright / PyTest) * AI Evaluation Metrics ...
Denver, MO · On-site +1
Prompt Engineering Validation * AI Model Validation & Evaluation * API Testing (Postman, REST APIs, Swagger) * Python * Test Automation (Selenium / Playwright / PyTest) * AI Evaluation Metrics ...
Design and own end-to-end LLM evaluation frameworks - including automated prompt regression ... Define the AI testing roadmap and quality standards for the program - establishing evaluation ...
Design and own end-to-end LLM evaluation frameworks - including automated prompt regression ... Define the AI testing roadmap and quality standards for the program - establishing evaluation ...
Jersey City, NJ · On-site
AI Content Designer; NLP Prompt Specialist Core keywords prompt engineering, system prompt, few ... Prompt testing, versioning, evaluation, and quality-improvement workflows. * Reusable prompt ...
Jersey City, NJ · On-site
AI Content Designer; NLP Prompt Specialist Core keywords prompt engineering, system prompt, few ... Prompt testing, versioning, evaluation, and quality-improvement workflows. * Reusable prompt ...
Develop RAG pipelines, prompt engineering, function/tool calling, context management, and agent ... Participate in architecture, development, testing, deployment, monitoring, and production support.
New
Develop RAG pipelines, prompt engineering, function/tool calling, context management, and agent ... Participate in architecture, development, testing, deployment, monitoring, and production support.
New
... versions and prompt variations. • Build comprehensive test suites for agentic AI systems ... testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and ...
... versions and prompt variations. • Build comprehensive test suites for agentic AI systems ... testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and ...
Denver, MO · Remote
Prompt Engineering Validation * AI Model Validation & Evaluation * API Testing (Postman, REST APIs, Swagger) * Python * Test Automation (Selenium / Playwright / PyTest) * AI Evaluation Metrics ...
Quick apply
Denver, MO · Remote
Prompt Engineering Validation * AI Model Validation & Evaluation * API Testing (Postman, REST APIs, Swagger) * Python * Test Automation (Selenium / Playwright / PyTest) * AI Evaluation Metrics ...
$14.18 - $14.90
4% of jobs
$14.90 - $15.63
13% of jobs
$16.18 is the 25th percentile. Wages below this are outliers.
$15.63 - $16.35
11% of jobs
$16.35 - $17.07
7% of jobs
$17.07 - $17.79
7% of jobs
The median wage is $18.24 / hr.
$17.79 - $18.51
13% of jobs
$18.51 - $19.23
13% of jobs
$19.75 is the 75th percentile. Wages above this are outliers.
$19.23 - $19.95
11% of jobs
$19.95 - $20.67
8% of jobs
$20.67 - $21.39
7% of jobs
$21.39 - $22.12
6% of jobs
$14
$18
$22
AI Prompt Testers often encounter challenges such as ensuring prompt accuracy across diverse contexts, identifying subtle biases or inconsistencies in AI-generated outputs, and adapting to frequent software updates or changes in model behavior. They may need to balance thorough testing with efficiency under tight deadlines, and clearly communicate feedback to developers or product teams. Tackling these challenges requires strong analytical skills, adaptability, and close collaboration with cross-functional teams to help improve product quality and user experience.
To thrive as an AI Prompt Tester, you need a strong grasp of language comprehension, analytical thinking, and familiarity with AI-driven content generation platforms. Experience using tools like GPT-based interfaces, prompt management systems, and possibly certifications in machine learning or natural language processing can be advantageous. Attention to detail, problem-solving, and effective communication are valuable soft skills in this role. These skills are crucial for ensuring high-quality, relevant AI outputs and facilitating continuous improvements in AI prompt performance.
An AI Prompt Tester evaluates and refines AI-generated responses by testing different prompts to ensure accuracy, relevance, and coherence. They analyze outputs, identify inconsistencies, and provide feedback to improve AI models. This role requires critical thinking, attention to detail, and a strong understanding of AI behavior. It is commonly used in AI development, chatbots, and content generation systems.
Cities with the most Ai Prompt Tester job openings:
States with the most job openings for Ai Prompt Tester jobs include:
The top searched job categories for Ai Prompt Tester jobs are:

Full-time
Re-posted 22 days ago
*Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time.*
Location: Dallas, TX (hybrid)
Job Description
We are seeking a Product Owner (AI & Prompt Strategy) to serve as a prompt-development focused leader for conversational AI initiatives. Operating within a Federated Hub-and-Spoke model, you will sit at the orchestration level across work teams to own the company-wide prompt and persona strategy. Partnering closely with the core engineering team and leading a team of Product Analysts, you will drive the parallel development and deployment of voice AI agents across CDL/Sales, Servicing, and Fulfillment.
Key Responsibilities:
· Strategy & Orchestration: Lead the overall product vision for conversational AI agents, owning the company-wide persona, brand voice, and overarching prompt strategy.
· Agent & Sub-Agent Architecture Design: Architect the multi-agent orchestration layer, defining how primary routing agents interpret user intent and seamlessly hand off tasks to specialized, domain-specific sub-agents (e.g., specific agents for Sales vs. Servicing vs. Fulfillment). Establish the logic boundaries, tool access, and context parameters for each sub-agent.
· Iterative Prompt Refinement: Continuously develop, test, and tune complex system prompts, system instructions, and few-shot examples. Optimize these prompts across the AWS tech stack to balance conversational quality, token efficiency, and response latency.
· Prompt Effectiveness Tracking & Analytics: Establish robust tracking frameworks to measure prompt and agent efficacy at a granular level. Monitor specialized metrics such as hallucination rates, contextual accuracy, task completion/containment rates, and user drop-off to drive continuous prompt optimization cycles.
· Stakeholder Alignment: Serve as the central orchestrator across CDL/Sales, Servicing, and Fulfillment teams to negotiate and refine business requirements, ensuring parallel development without technical bottlenecks.
· Testing & Quality Assurance: Define the "Golden Dataset" of test scenarios to feed into engineering’s automated LLM-as-a-judge evaluation pipelines. Oversee structured A/B testing initiatives and data analysis to iteratively optimize agent responses against baseline configurations.
· Performance & ROI: Define product strategy and track ROI by aligning AI agent performance to key business KPIs, including Customer Satisfaction (CSAT) and Average Handle Time (AHT).
· Cross-Functional Collaboration: Partner with the engineering team to integrate prompt designs with the core infrastructure, latency management pipelines, and real-time safety guardrails.
What You’ll Bring
· 5–7 years of Product Management or Product Ownership experience, ideally within mortgage, fintech, or complex customer servicing environments.
· Demonstrated experience with AI/ML products, specifically conversational AI, Large Language Models (LLMs), and deep prompt engineering strategies within an enterprise AWS environment.
· Proven ability to design multi-agent or complex state-machine conversational architectures.
· Strong analytical background with a track record of building effectiveness dashboards, running A/B tests, and utilizing AI-specific metrics (e.g., hallucination tracking, prompt latency) to drive improvements.
· Ability to translate complex business logic into precise technical configurations and natural language flows.
· Exceptional leadership and stakeholder management skills.
Kaleidoscope, an Infosys Company, is an equal opportunity employer, and all qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, spouse of protected veteran, or disability.