1

Ai Evaluation Engineer Jobs (NOW HIRING)

Zof AI is seeking an AI Evaluation Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the ...

Zof AI is seeking an AI Evaluation Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the ...

Deposco is seeking an experienced AI Evaluation Engineer to join our innovative quality assurance team. This role focuses on ensuring the accuracy, reliability, and performance of AI-driven ...

AI Evaluation Engineer

Washington, DC ยท On-site

$81K - $105K/yr

Ensure enterprise AI platforms and solutions are reliable, tested, and production-ready. Remote / Hybrid Security, Quality & Reliability Full-time Role Overview The QA Engineer is responsible for ...

Ciklum is looking for a Senior AI Evaluation Engineer to join our team full-time in the US . We are a custom product engineering company that supports both multinational organizations and scaling ...

Machine Learning Engineer - Agentic AI Evaluation Frameworks Cupertino, California, United States Machine Learning and AI Imagine what you could do here. At Apple, great ideas have a way of becoming ...

New

AI Engineer, Evaluation

New York, NY ยท On-site

$150K - $250K/yr

AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines ...

Manager; AI Evaluation Engineering

Chicago, IL ยท On-site

$147K - $240K/yr

Join the AI Engineering team of Cat Digital and take charge of leading a team dedicated to evaluating and validating our advanced generative AI solutions-including intelligent agents, digital ...

next page

Showing results 1-20

Ai Evaluation Engineer information

See salary details

$25

$53

$76

How much do ai evaluation engineer jobs pay per hour?

As of Sep 14, 2026, the average hourly pay for ai evaluation engineer in the United States is $53.63, according to ZipRecruiter salary data. Most workers in this role earn between $43.27 and $62.26 per hour, depending on experience, location, and employer.

How to become an AI evaluation engineer?

To become an AI evaluation engineer, candidates typically need a bachelor's or master's degree in computer science, data science, or a related field. Strong skills in machine learning, programming (Python, R), and understanding of AI models are essential, along with experience in data analysis and evaluation metrics. Gaining familiarity with AI frameworks and tools, as well as relevant certifications, can enhance job prospects.

What is the role of AI evaluation engineer?

An AI evaluation engineer is responsible for assessing the performance, accuracy, and fairness of artificial intelligence models. They develop testing protocols, analyze model outputs, and ensure AI systems meet quality and ethical standards, often using tools like benchmarking datasets and evaluation metrics. This role requires strong analytical skills and knowledge of machine learning frameworks.
More about Ai Evaluation Engineer jobs

What cities are hiring for Ai Evaluation Engineer jobs?

Cities with the most Ai Evaluation Engineer job openings:

What states have the most Ai Evaluation Engineer jobs?

States with the most job openings for Ai Evaluation Engineer jobs include:

What are popular job titles related to Ai Evaluation Engineer jobs?

For Ai Evaluation Engineer jobs, the most frequently searched job titles are:

Infographic showing various Ai Evaluation Engineer job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 76% Full Time, 20% Part Time, and 3% Contract. Highlights an 62% Physical, 4% Hybrid, and 34% Remote job distribution, with an average salary of $111,552 per year, or $53.6 per hour.

AI Evaluation Engineer

San Francisco, CA โ€ข On-site

Other

Posted 15 days ago


Job description

Zof AI is seeking an AI Evaluation Engineer to build the tests that determine whether AI actually works. Verification is the heart of what Zof AI does, so this role sits close to the core of the product: designing eval suites, verification harnesses, and quality gates that confirm AI systems built the right thing, part QA discipline and part domain judgment. The ideal candidate is skeptical by default, rigorous about measurement, and motivated by turning "it seems to work" into evidence.

Engineering ยท Mid to Senior ยท Full-time ยท On-site ยท San Francisco, CA

Responsibilities
  • Design and build eval suites for AI products and agent systems.
  • Build verification harnesses that confirm the AI built the right thing.
  • Define quality gates that gate what ships and what does not.
  • Turn domain expertise and customer requirements into testable checks.
  • Hunt failure modes: regressions, hallucinations, and silent errors.
  • Make eval results legible to engineers, product, and customers.
  • Wire evals into CI and the development loop.
  • Raise the standard for what "working" means across the company.
Requirements
  • Experience testing, evaluating, or QA-ing complex software systems.
  • Understanding of how LLM and agent systems fail.
  • Strong analytical rigor and skepticism.
  • Ability to write code to build harnesses and automation.
  • Attention to detail and a high quality bar.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership.
Nice to have
  • Experience building LLM evals, benchmarks, or test infrastructure.
  • QA, SDET, or test automation background.
  • Domain expertise in a vertical where correctness matters.
  • Experience with statistical evaluation methods.

Must understand how to measure whether AI systems actually work, beyond demos

#J-18808-Ljbffr