Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one ... Demonstrated expertise testing LLM and Gen AI systems - including prompt testing, output evaluation ...
Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one ... Demonstrated expertise testing LLM and Gen AI systems - including prompt testing, output evaluation ...
... day one - reviewing prompt architectures, agent designs, and system workflows for testability and ... expertise testing LLM and Gen AI systems - including prompt testing, output evaluation ...
... day one - reviewing prompt architectures, agent designs, and system workflows for testability and ... expertise testing LLM and Gen AI systems - including prompt testing, output evaluation ...
Scrum Tester with Agent AI
Hartford, CT · On-site
Hartford, CT (Onsite from Day 1) Job Type: Contract Skill Metrics: AI Testing Jira Java Selenium Top skills required for this role: 1. Agent AI - Prompt Generation 2. Selenium Automation 3. AI tools ...
Quick apply
Scrum Tester with Agent AI
Hartford, CT · On-site
Hartford, CT (Onsite from Day 1) Job Type: Contract Skill Metrics: AI Testing Jira Java Selenium Top skills required for this role: 1. Agent AI - Prompt Generation 2. Selenium Automation 3. AI tools ...
Manager, AI Quality & Reliability Engineering
Edina, MN · On-site
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
Edina, MN · On-site
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
Chicago, IL · On-site
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
Chicago, IL · On-site
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
Irving, TX · On-site
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
Irving, TX · On-site
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
Edina, MN · On-site
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
Edina, MN · On-site
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Manager, AI Quality & Reliability Engineering
$88K - $155K/yr
This role combines hands-on quality engineering leadership, AI-enabled testing modernization, healt ... Responsibilities * Lead day-to-day AI Quality Engineering activities supporting AI-powered ...
Lead AI Compliance Testing
Frisco, TX · Hybrid
$146K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
Frisco, TX · Hybrid
$146K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
Columbus, OH · Hybrid
$151K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
Columbus, OH · Hybrid
$151K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
Columbus, OH · On-site
$151K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
Columbus, OH · On-site
$151K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
New York, NY · Hybrid
$171K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
New York, NY · Hybrid
$171K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
Draper, UT · Hybrid
$146K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
Draper, UT · Hybrid
$146K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
Chadds Ford, PA · Hybrid
$155K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Lead AI Compliance Testing
Chadds Ford, PA · Hybrid
$155K/yr
AI Testing & Credible Challenge: Execute the quarterly, risk-based testing program to perform ... Normal office environment. (Remote or Hybrid), 3 to 4 days per month are required in office if ...
Manager, Programs & AI Delivery
Needham, MA · On-site
Own the day-to-day AI build lifecycle - from problem framing and prototyping through testing, deployment, and iteration. * Document how tools work and apply sensible quality and evaluation checks so ...
Manager, Programs & AI Delivery
Needham, MA · On-site
Own the day-to-day AI build lifecycle - from problem framing and prototyping through testing, deployment, and iteration. * Document how tools work and apply sensible quality and evaluation checks so ...
AI/ML Engineer
Englewood, CO · On-site
Serve as primary day-to-day AI/ML expert & performer other team members are AI/ML learners or ... Develop and oversee robust validation and testing frameworks to ensure that AI/ML models meet ...
AI/ML Engineer
Englewood, CO · On-site
Serve as primary day-to-day AI/ML expert & performer other team members are AI/ML learners or ... Develop and oversee robust validation and testing frameworks to ensure that AI/ML models meet ...
QA Tester - OSS AI Testing
Manhattan, NY · On-site
Agile Framework Experience, JIRA thoroughness * 5 Days in person at office (for NJ based talents) & rest 50% can be remote Nice To have: * Agentic AI Solutioning (or) relevant knowledge to leverage ...
QA Tester - OSS AI Testing
Manhattan, NY · On-site
Agile Framework Experience, JIRA thoroughness * 5 Days in person at office (for NJ based talents) & rest 50% can be remote Nice To have: * Agentic AI Solutioning (or) relevant knowledge to leverage ...
AI Lead Engineer
Nashville, TN · On-site
$99K - $130K/yr
... day. AI Lead Engineer Responsibilities: * Understand the define technical vision, roadmap, and ... testing automation, performance analysis, and operational intelligence. * Establish and enforce ...
AI Lead Engineer
Nashville, TN · On-site
$99K - $130K/yr
... day. AI Lead Engineer Responsibilities: * Understand the define technical vision, roadmap, and ... testing automation, performance analysis, and operational intelligence. * Establish and enforce ...
Day Ai Tester information
See salary details
$10.82 - $15.54
7% of jobs
$15.54 - $20.26
16% of jobs
$21.31 is the 25th percentile. Wages below this are outliers.
$20.26 - $24.98
9% of jobs
$24.98 - $29.70
3% of jobs
$29.70 - $34.42
10% of jobs
The median wage is $36.31 / hr.
$34.42 - $39.14
10% of jobs
$39.14 - $43.86
7% of jobs
$43.86 - $48.58
9% of jobs
$49.21 is the 75th percentile. Wages above this are outliers.
$48.58 - $53.30
16% of jobs
$53.30 - $58.02
6% of jobs
$58.02 - $62.74
5% of jobs
$10
$38
$62
How much do day ai tester jobs pay per hour?
How much do AI testers get paid?
What is the difference between Day Ai Tester vs Data Analyst?
| Aspect | Day Ai Tester | Data Analyst |
|---|---|---|
| Required Credentials | Basic knowledge of AI tools, testing certifications | Bachelor's in Data Science, Statistics, or related fields |
| Work Environment | Tech companies, AI development teams, testing labs | Business, finance, healthcare sectors analyzing data |
| Employer & Industry Usage | AI startups, tech firms, software companies | Corporations, consulting firms, research institutions |
| Common Search & Comparison | Yes | Yes |
The main difference between a Day Ai Tester and a Data Analyst lies in their focus. Day Ai Testers primarily evaluate AI systems for accuracy and functionality, often requiring knowledge of AI tools and testing certifications. Data Analysts interpret data to inform business decisions, typically holding degrees in data-related fields. While both roles work with data and are found in tech-driven industries, their daily tasks and skill requirements differ significantly.
Can I get paid to test AI?
What is a $900000 AI job?
How does a Day AI Tester typically collaborate with developers and data scientists during the model evaluation process?
How do I become an AI tester?
What are the key skills and qualifications needed to thrive as a Day AI Tester, and why are they important?
What are Day AI Testers?

Full-time
Medical, Dental, Vision, Life, Retirement, PTO
Posted 7 days ago
Job description
Our Purpose
Mastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we're helping build asustainableeconomy where everyone can prosper. We support a wide range of digital payments choices, making transactionssecure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.
Title and Summary
Manager, AI Engineering (Tester )Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for a AI Tester for the Operational Intelligence Program within B&MI. This is a highly specialized, hands-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise-ready. This role will lead AI quality engineering efforts - defining evaluation frameworks, red-teaming strategies, and LLMOps quality gates - while fostering a culture of rigorous, first-class AI testing across the program.Roles and Responsibilities:
Design and own end-to-end LLM evaluation frameworks - including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations.
Build comprehensive test suites for agentic AI systems - validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling across multi-step reasoning workflows.
Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, TruLens, and DeepEval.
Lead structured red-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation - building and maintaining an evolving adversarial test library.
Execute fairness, bias, and Responsible AI audits - testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy.
Design and run inference performance benchmarks - measuring latency, throughput, token efficiency, and degradation under peak load - and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS).
Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, CloudWatch).
Define the AI testing roadmap and quality standards for the program - establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI workstreams.
Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one - reviewing prompt architectures, agent designs, and system workflows for testability and risk.
Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, TruthfulQA, MT-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge.
All About You:
Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands-on experience leading AI/ML quality engineering or LLM testing programs in production environments.
Demonstrated expertise testing LLM and Gen AI systems - including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings.
Deep hands-on knowledge of AI evaluation frameworks and tooling: RAGAS, DeepEval, TruLens, LangSmith, PromptFlow, Weights & Biases Evals, or equivalent platforms.
Strong understanding of Gen AI failure modes - hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures - and proven methods to surface and document them systematically.
Strong Python programming skills with the ability to independently build test automation scripts, evaluation pipelines, and API-level integration tests; SQL proficiency required.
Working knowledge of LLM ecosystems - OpenAI, Anthropic, Hugging Face, LangChain/LangGraph - sufficient to understand model behavior, prompt structure, and agent architecture deeply enough to test them rigorously.
Familiarity with MLOps/LLMOps pipelines (MLflow, Databricks, SageMaker) and experience integrating automated quality gates into CI/CD workflows for AI systems.
Experience with cloud AI infrastructure (AWS, Azure, or GCP) and observability tooling for monitoring live AI system behavior and output quality in production.
Strong analytical, communication, and stakeholder management skills - with the ability to translate complex AI failure patterns into clear risk assessments and remediation recommendations for both technical and business audiences.Mastercard is a merit-based, inclusive, equal opportunity employer that considers applicants without regard to gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. We hire the most qualified candidate for the role. In the US or Canada, if you require accommodations or assistance to complete the online application process or during the recruitment process, please contact reasonable_accommodation@mastercard.com and identify the type of accommodation or assistance you are requesting. Do not include any medical or health information in this email. The Reasonable Accommodations team will respond to your email promptly.
Corporate Security Responsibility
All activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:
Abide by Mastercard's security policies and practices;
Ensure the confidentiality and integrity of the information being accessed;
Report any suspected information security violation or breach, and
Complete all periodic mandatory security trainings in accordance with Mastercard's guidelines.
Pay Ranges
O'Fallon, Missouri: $140,000 - $231,000 USD