Title and Summary Manager, AI Engineering (Tester ) Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help ...
Title and Summary Manager, AI Engineering (Tester ) Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help ...
Title and Summary Manager, AI Engineering (Tester ) Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help ...
Title and Summary Manager, AI Engineering (Tester ) Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help ...
AI Agentic Tester
Denver, MO · On-site +1
AI Agentic Tester Denver, MO (Remote) Must-Have Skills AI Agentic Testing Functional Testing Generative AI Testing Large Language Models (LLMs) AI Agents & Agentic AI Retrieval-Augmented Generation ...
AI Agentic Tester
Denver, MO · On-site +1
AI Agentic Tester Denver, MO (Remote) Must-Have Skills AI Agentic Testing Functional Testing Generative AI Testing Large Language Models (LLMs) AI Agents & Agentic AI Retrieval-Augmented Generation ...
AI Agentic Tester
Denver, MO · On-site +1
AI Agentic Tester Denver, MO (Remote) Must-Have Skills * AI Agentic Testing * Functional Testing * Generative AI Testing * Large Language Models (LLMs) * AI Agents & Agentic AI * Retrieval-Augmented ...
AI Agentic Tester
Denver, MO · On-site +1
AI Agentic Tester Denver, MO (Remote) Must-Have Skills * AI Agentic Testing * Functional Testing * Generative AI Testing * Large Language Models (LLMs) * AI Agents & Agentic AI * Retrieval-Augmented ...
AI Agentic Tester (Remote)
Denver, MO · Remote
AI Agentic Tester Denver, MO (Remote) Must-Have Skills * AI Agentic Testing * Functional Testing * Generative AI Testing * Large Language Models (LLMs) * AI Agents & Agentic AI * Retrieval-Augmented ...
Quick apply
AI Agentic Tester (Remote)
Denver, MO · Remote
AI Agentic Tester Denver, MO (Remote) Must-Have Skills * AI Agentic Testing * Functional Testing * Generative AI Testing * Large Language Models (LLMs) * AI Agents & Agentic AI * Retrieval-Augmented ...
The ideal candidate will lead system integration testing, ensure adherence to quality standards, and develop AI-specific automated testing frameworks. Responsibilities : • Conduct system ...
The ideal candidate will lead system integration testing, ensure adherence to quality standards, and develop AI-specific automated testing frameworks. Responsibilities : • Conduct system ...
The ideal candidate will lead system integration testing, ensure adherence to quality standards, and develop AI-specific automated testing frameworks. Responsibilities : • Conduct system ...
The ideal candidate will lead system integration testing, ensure adherence to quality standards, and develop AI-specific automated testing frameworks. Responsibilities : • Conduct system ...
The ideal candidate will lead system integration testing, ensure adherence to quality standards, and develop AI-specific automated testing frameworks. Responsibilities : • Conduct system ...
The ideal candidate will lead system integration testing, ensure adherence to quality standards, and develop AI-specific automated testing frameworks. Responsibilities : • Conduct system ...
Create the future of AI testing: Design and develop new AI adversarial testing methodologies that become part of BMO's long-term security testing strategy. * Expand your expertise: Lead security ...
Create the future of AI testing: Design and develop new AI adversarial testing methodologies that become part of BMO's long-term security testing strategy. * Expand your expertise: Lead security ...
Scrum Tester with Agent AI
Hartford, CT · On-site
Scrum Tester with Agent AI Location: Hartford, CT (Onsite from Day 1) Job Type: Contract Skill Metrics: AI Testing Jira Java Selenium Top skills required for this role: 1. Agent AI - Prompt ...
Quick apply
Scrum Tester with Agent AI
Hartford, CT · On-site
Scrum Tester with Agent AI Location: Hartford, CT (Onsite from Day 1) Job Type: Contract Skill Metrics: AI Testing Jira Java Selenium Top skills required for this role: 1. Agent AI - Prompt ...
They are seeking a highly motivated QA Tester with strong experience in both manual and automation testing, specifically with Playwright and AI-driven testing tools. The role involves ensuring high ...
They are seeking a highly motivated QA Tester with strong experience in both manual and automation testing, specifically with Playwright and AI-driven testing tools. The role involves ensuring high ...
QA Tester
Fremont, CA · On-site
AI awareness and understanding of AI-powered applications and tools. * Strong analytical and ... Experience testing AI-enabled or GenAI applications. * Exposure to Agentic AI testing concepts and ...
QA Tester
Fremont, CA · On-site
AI awareness and understanding of AI-powered applications and tools. * Strong analytical and ... Experience testing AI-enabled or GenAI applications. * Exposure to Agentic AI testing concepts and ...
AI Testing Specialist
Manhattan, NY · On-site
Diverse Lynx is seeking an innovative AI Testing Specialist to lead the adoption of AI-powered Quality Engineering practices across the Software Development Life Cycle. The role involves designing ...
AI Testing Specialist
Manhattan, NY · On-site
Diverse Lynx is seeking an innovative AI Testing Specialist to lead the adoption of AI-powered Quality Engineering practices across the Software Development Life Cycle. The role involves designing ...
AI Developer - ServiceNow
Cincinnati, OH · Remote
$55.25 - $76/hr
AI Testing & Quality Practices * AI Agent Design & Enhancement * Technical Design Agents * Configuration (Config) Agents * AI Testing & Quality-Focused Agents * Prompt Engineering & Optimization
Quick apply
AI Developer - ServiceNow
Cincinnati, OH · Remote
$55.25 - $76/hr
AI Testing & Quality Practices * AI Agent Design & Enhancement * Technical Design Agents * Configuration (Config) Agents * AI Testing & Quality-Focused Agents * Prompt Engineering & Optimization
AI Testing Specialist
New York, NY · On-site
$100K - $115K/yr
AI Testing Specialist The Testing Specialist is a forward-looking Quality Engineering role centered on integrating artificial intelligence and automated validation into every phase of the Software ...
AI Testing Specialist
New York, NY · On-site
$100K - $115K/yr
AI Testing Specialist The Testing Specialist is a forward-looking Quality Engineering role centered on integrating artificial intelligence and automated validation into every phase of the Software ...
QA Tester
Fremont, CA · On-site
Good to Have • Experience testing AI-enabled or GenAI applications. • Exposure to Agentic AI testing concepts and approaches. • Experience with API testing tools such as Postman or Rest Assured ...
QA Tester
Fremont, CA · On-site
Good to Have • Experience testing AI-enabled or GenAI applications. • Exposure to Agentic AI testing concepts and approaches. • Experience with API testing tools such as Postman or Rest Assured ...
QA Automation Tester
Chicago, IL · On-site
San Francisco, CA/ Chicago, IL • 60% is Device Testing and 40% is LLM / GenAI Validation (AI Testing) • Test automation experience is non-negotiable - manual-only profiles will not be considered ...
QA Automation Tester
Chicago, IL · On-site
San Francisco, CA/ Chicago, IL • 60% is Device Testing and 40% is LLM / GenAI Validation (AI Testing) • Test automation experience is non-negotiable - manual-only profiles will not be considered ...
AI Testing Specialist
Manhattan, NY · On-site
Tata Consultancy Services is seeking an AI Testing Specialist, a forward-looking Quality Engineering role focused on integrating artificial intelligence into the Software Development Life Cycle. The ...
AI Testing Specialist
Manhattan, NY · On-site
Tata Consultancy Services is seeking an AI Testing Specialist, a forward-looking Quality Engineering role focused on integrating artificial intelligence into the Software Development Life Cycle. The ...
Senior AI Test Engineer (Remote)
Sterling, VA · On-site +1
$50 - $80/hr
OWASP AI Testing Guide, NIST AI Risk Management Framework (AI RMF), NIST Cybersecurity Framework (CSF) Cyber AI Profile, AI Bills of Materials (AI-BOMs), NSA AI Security Guidelines. Location: Remote ...
Senior AI Test Engineer (Remote)
Sterling, VA · On-site +1
$50 - $80/hr
OWASP AI Testing Guide, NIST AI Risk Management Framework (AI RMF), NIST Cybersecurity Framework (CSF) Cyber AI Profile, AI Bills of Materials (AI-BOMs), NSA AI Security Guidelines. Location: Remote ...
OWASP AI Testing Guide, NIST AI Risk Management Framework (AI RMF), NIST Cybersecurity Framework (CSF) Cyber AI Profile, AI Bills of Materials (AI-BOMs), NSA AI Security Guidelines. Location: Remote ...
Quick apply
OWASP AI Testing Guide, NIST AI Risk Management Framework (AI RMF), NIST Cybersecurity Framework (CSF) Cyber AI Profile, AI Bills of Materials (AI-BOMs), NSA AI Security Guidelines. Location: Remote ...
Ai Tester information
See salary details
$10.82 - $15.54
7% of jobs
$15.54 - $20.26
16% of jobs
$21.31 is the 25th percentile. Wages below this are outliers.
$20.26 - $24.98
9% of jobs
$24.98 - $29.70
3% of jobs
$29.70 - $34.42
10% of jobs
The median wage is $36.31 / hr.
$34.42 - $39.14
10% of jobs
$39.14 - $43.86
7% of jobs
$43.86 - $48.58
9% of jobs
$49.21 is the 75th percentile. Wages above this are outliers.
$48.58 - $53.30
16% of jobs
$53.30 - $58.02
6% of jobs
$58.02 - $62.74
5% of jobs
$10
$38
$62
How much do ai tester jobs pay per hour?
What is an AI tester?
An AI Tester is responsible for evaluating artificial intelligence systems to ensure they function correctly, efficiently, and ethically. They design and execute test cases, identify flaws or biases, and verify that AI models meet performance standards. AI Testers work with developers and data scientists to improve AI reliability and user experience. Their role is crucial in preventing errors, reducing risks, and ensuring AI models make accurate and fair decisions.
What are the key skills and qualifications needed to thrive as an AI tester?
To thrive as an AI Tester, you need a background in computer science, experience with software testing methodologies, and a solid understanding of artificial intelligence technologies. Familiarity with testing tools (such as Selenium, Jupyter notebooks, or TensorFlow testing frameworks), programming languages like Python, and relevant certifications (e.g., ISTQB) are highly advantageous. Attention to detail, problem-solving abilities, and strong communication skills help AI Testers identify and articulate issues effectively. These skills ensure AI systems are reliable, accurate, and deliver expected outcomes in real-world applications.
What are some typical challenges faced by AI testers in their daily work?
AI Testers often encounter challenges such as managing and testing large, complex datasets, handling rapidly evolving algorithms, and ensuring consistent test coverage across various real-world scenarios. They must also validate that AI models are free from bias and produce accurate, reproducible results under different conditions. Overcoming these challenges requires both technical proficiency and adaptability. Collaboration with data scientists, developers, and product managers is common, and testers frequently update their testing approaches to keep pace with the fast changes in AI technology.
Can I get an AI Tester job with no experience?
How do you become an AI tester?
How much do AI testers make?
What cities are hiring for Ai Tester jobs?
Cities with the most Ai Tester job openings:
What are the most commonly searched types of Ai Tester jobs?
The most popular types of Ai Tester jobs are:
What states have the most Ai Tester jobs?
States with the most job openings for Ai Tester jobs include:
What job categories do people searching Ai Tester jobs look for?
The top searched job categories for Ai Tester jobs are:

Full-time
Medical, Dental, Vision, Life, Retirement, PTO
Re-posted 11 days ago
Key responsibilities
Design and own end-to-end LLM evaluation frameworks, including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection.
Build comprehensive test suites for agentic AI systems, validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling.
Lead structured red-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation.
Job description
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build asustainableeconomy where everyone can prosper. We support a wide range of digital payments choices, making transactionssecure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Manager, AI Engineering (Tester )Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for a AI Tester for the Operational Intelligence Program within B&MI. This is a highly specialized, hands-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise-ready. This role will lead AI quality engineering efforts - defining evaluation frameworks, red-teaming strategies, and LLMOps quality gates - while fostering a culture of rigorous, first-class AI testing across the program.Roles and Responsibilities:
Design and own end-to-end LLM evaluation frameworks - including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations.
Build comprehensive test suites for agentic AI systems - validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling across multi-step reasoning workflows.
Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, TruLens, and DeepEval.
Lead structured red-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation - building and maintaining an evolving adversarial test library.
Execute fairness, bias, and Responsible AI audits - testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy.
Design and run inference performance benchmarks - measuring latency, throughput, token efficiency, and degradation under peak load - and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS).
Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, CloudWatch).
Define the AI testing roadmap and quality standards for the program - establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI workstreams.
Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one - reviewing prompt architectures, agent designs, and system workflows for testability and risk.
Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, TruthfulQA, MT-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge.
All About You:
Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands-on experience leading AI/ML quality engineering or LLM testing programs in production environments.
Demonstrated expertise testing LLM and Gen AI systems - including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings.
Deep hands-on knowledge of AI evaluation frameworks and tooling: RAGAS, DeepEval, TruLens, LangSmith, PromptFlow, Weights & Biases Evals, or equivalent platforms.
Strong understanding of Gen AI failure modes - hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures - and proven methods to surface and document them systematically.
Strong Python programming skills with the ability to independently build test automation scripts, evaluation pipelines, and API-level integration tests; SQL proficiency required.
Working knowledge of LLM ecosystems - OpenAI, Anthropic, Hugging Face, LangChain/LangGraph - sufficient to understand model behavior, prompt structure, and agent architecture deeply enough to test them rigorously.
Familiarity with MLOps/LLMOps pipelines (MLflow, Databricks, SageMaker) and experience integrating automated quality gates into CI/CD workflows for AI systems.
Experience with cloud AI infrastructure (AWS, Azure, or GCP) and observability tooling for monitoring live AI system behavior and output quality in production.
Strong analytical, communication, and stakeholder management skills - with the ability to translate complex AI failure patterns into clear risk assessments and remediation recommendations for both technical and business audiences.Mastercard is a merit-based, inclusive, equal opportunity employer that considers applicants without regard to gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. We hire the most qualified candidate for the role. In the US or Canada, if you require accommodations or assistance to complete the online application process or during the recruitment process, please contact reasonable_accommodation@mastercard.com and identify the type of accommodation or assistance you are requesting. Do not include any medical or health information in this email. The Reasonable Accommodations team will respond to your email promptly.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
Abide by Mastercard's security policies and practices;
Ensure the confidentiality and integrity of the information being accessed;
Report any suspected information security violation or breach, and
Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Pay Ranges
O'Fallon, Missouri: $140,000 - $231,000 USD