What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Scientist
Mclean, VA · On-site
$105 - $145/hr
The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...
AI Evaluation Scientist
Mclean, VA · On-site
$105 - $145/hr
The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
AI Evaluation Engineer (Python, QA or Security)
Queens, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Queens, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Austin, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Austin, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Houston, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Houston, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Houston, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Houston, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
New York, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
New York, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$150 - $230/hr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$150 - $230/hr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$147K - $232K/yr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$147K - $232K/yr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Lead Forward Deployed Engineer, AI Evaluation Platform
Seattle, WA · On-site
$175 - $263.30/hr
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are building the scientific foundation and self-service tools for how AI evaluation is done at scale ...
Lead Forward Deployed Engineer, AI Evaluation Platform
Seattle, WA · On-site
$175 - $263.30/hr
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are building the scientific foundation and self-service tools for how AI evaluation is done at scale ...
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are building the scientific foundation and self-service tools for how AI evaluation is done at scale ...
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are building the scientific foundation and self-service tools for how AI evaluation is done at scale ...
Sr. Evaluation Engineer
San Francisco, CA · On-site
$123K - $169K/yr
What You'll Do: LogicMonitor is the AI-first hybrid observability platform powering the next ... As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide ...
New
Sr. Evaluation Engineer
San Francisco, CA · On-site
$123K - $169K/yr
What You'll Do: LogicMonitor is the AI-first hybrid observability platform powering the next ... As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide ...
New
Sr. Evaluation Engineer
San Francisco, CA · On-site
$123K - $169K/yr
As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation ...
New
Sr. Evaluation Engineer
San Francisco, CA · On-site
$123K - $169K/yr
As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation ...
New
AIML - Software Engineer - AI, Evaluation
$150K - $277K/yr
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
AIML - Software Engineer - AI, Evaluation
$150K - $277K/yr
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · Remote
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · Remote
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · On-site +1
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · On-site +1
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Quick apply
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Ai Evaluation Engineer information
See salary details
$25.48 - $30.14
1% of jobs
$30.14 - $34.79
5% of jobs
$34.79 - $39.44
9% of jobs
$43.46 is the 25th percentile. Wages below this are outliers.
$39.44 - $44.10
12% of jobs
$44.10 - $48.75
10% of jobs
The median wage is $53.08 / hr.
$48.75 - $53.41
15% of jobs
$53.41 - $58.06
15% of jobs
$61.36 is the 75th percentile. Wages above this are outliers.
$58.06 - $62.72
13% of jobs
$62.72 - $67.37
10% of jobs
$67.37 - $72.03
10% of jobs
$72.03 - $76.68
2% of jobs
$25
$53
$76
How much do ai evaluation engineer jobs pay per hour?
How to become an AI evaluation engineer?
What is the role of AI evaluation engineer?
What cities are hiring for Ai Evaluation Engineer jobs?
Cities with the most Ai Evaluation Engineer job openings:
What states have the most Ai Evaluation Engineer jobs?
States with the most job openings for Ai Evaluation Engineer jobs include:
What job categories do people searching Ai Evaluation Engineer jobs look for?
The top searched job categories for Ai Evaluation Engineer jobs are:

$50/hr
Part-time
Posted 17 days ago
Job description
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.
You'll create challenging tasks and evaluation criteria within realistic simulated environments:
- Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
- Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
- Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
- Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust
What this is NOT
- Not data labeling
- Not prompt engineering
- Not writing code from scratch - the agent writes most of the code; you guide and evaluate
What we look for
- 5+ years in software development
- Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
- Experience writing tests (functional, integration)
- English proficiency - B2+
Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
How it works
Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid
Compensation
Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.