1

Ai Evaluation Engineer Jobs (NOW HIRING)

AI Evaluation Scientist

Mclean, VA · On-site

$105 - $145/hr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...

... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...

Sr. Evaluation Engineer

San Francisco, CA · On-site

$123K - $169K/yr

What You'll Do: LogicMonitor is the AI-first hybrid observability platform powering the next ... As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide ...

New

Sr. Evaluation Engineer

San Francisco, CA · On-site

$123K - $169K/yr

As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation ...

New

Showing results 41-60

Ai Evaluation Engineer information

See salary details

$25

$53

$76

How much do ai evaluation engineer jobs pay per hour?

As of Aug 22, 2026, the average hourly pay for ai evaluation engineer in the United States is $53.63, according to ZipRecruiter salary data. Most workers in this role earn between $43.27 and $62.26 per hour, depending on experience, location, and employer.

How to become an AI evaluation engineer?

To become an AI evaluation engineer, candidates typically need a bachelor's or master's degree in computer science, data science, or a related field. Strong skills in machine learning, programming (Python, R), and understanding of AI models are essential, along with experience in data analysis and evaluation metrics. Gaining familiarity with AI frameworks and tools, as well as relevant certifications, can enhance job prospects.

What is the role of AI evaluation engineer?

An AI evaluation engineer is responsible for assessing the performance, accuracy, and fairness of artificial intelligence models. They develop testing protocols, analyze model outputs, and ensure AI systems meet quality and ethical standards, often using tools like benchmarking datasets and evaluation metrics. This role requires strong analytical skills and knowledge of machine learning frameworks.
More about Ai Evaluation Engineer jobs

What cities are hiring for Ai Evaluation Engineer jobs?

Cities with the most Ai Evaluation Engineer job openings:

What states have the most Ai Evaluation Engineer jobs?

States with the most job openings for Ai Evaluation Engineer jobs include:

Infographic showing various Ai Evaluation Engineer job openings in the United States as of August 2026, with employment types broken down into 76% Full Time, 21% Part Time, and 3% Contract. Highlights an 64% Physical, 4% Hybrid, and 32% Remote job distribution, with an average salary of $111,552 per year, or $53.6 per hour.

AI Evaluation Engineer (Python, QA or Security)

Mindrift

Remote

$50/hr

Part-time

Posted 17 days ago


Job description

Please submit your CV in English and indicate your level of English proficiency.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.
You'll create challenging tasks and evaluation criteria within realistic simulated environments:
  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what "solved" means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

What this is NOT
  • Not data labeling
  • Not prompt engineering
  • Not writing code from scratch - the agent writes most of the code; you guide and evaluate

What we look for
  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Why this is hard
Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.
How it works
Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid
Compensation
Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.