1

Ai Evaluation Jobs (NOW HIRING)

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

Highlight both strengths and areas for improvement in AI-generated outputs. * Apply consistent judgment across a wide range of evaluation tasks. Maintain Evaluation Quality * Follow detailed project ...

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission ...

AI Evaluation Scientist

Mclean, VA · On-site

$105K - $145K/yr

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission ...

AI Evaluation Scientist

Mclean, VA · On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA · On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA · On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission ...

Requirements Key Responsibilities Design AI Evaluation Tasks * Create realistic, multi-step benchmark tasks based on professional technology workflows. * Develop challenges using technical ...

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI/ML Evaluation Engineer

Atlanta, GA · On-site

$79K - $105K/yr

R0243409 AI/ML Evaluation Engineer The Opportunity: As an experience d engineer, you know that machine learning ( ML ) and AI evaluation are critical to understanding and operationalizing massive ...

AI Evaluation Engineer

Manhattan, NY · On-site

$161K - $200K/yr

Position Summary As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and ...

Posted today

Your evaluations and structured feedback will directly contribute to improving AI systems designed to support legal analysis and professional decision-making. This is a fully remote, independent ...

AI Evaluation Engineer

Denver, CO · On-site

$161K - $200K/yr

Position Summary As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and ...

Posted today

next page

Showing results 1-20

Ai Evaluation information

See salary details

$9

$17

$28

How much do ai evaluation jobs pay per hour?

As of Jul 19, 2026, the average hourly pay for ai evaluation in the United States is $17.85, according to ZipRecruiter salary data. Most workers in this role earn between $16.11 and $18.27 per hour, depending on experience, location, and employer.

What are some common challenges faced by professionals working in AI evaluation roles?

Professionals in AI evaluation often encounter challenges such as ensuring unbiased and accurate assessment of AI models, keeping up with rapidly evolving technologies, and interpreting complex data outputs. Balancing thoroughness with tight project deadlines can be demanding, as can collaborating with cross-functional teams like data scientists, product managers, and engineers to align evaluation metrics with business goals. Additionally, understanding and applying ethical guidelines in AI evaluation is increasingly important as organizations prioritize responsible AI deployment.

What are the key skills and qualifications needed to thrive as an AI Evaluator, and why are they important?

To thrive as an AI Evaluator, you need a strong analytical mindset, attention to detail, and familiarity with AI concepts, often supported by a relevant degree in computer science, linguistics, or a related field. Experience with data annotation platforms, evaluation tools, and sometimes knowledge of programming languages like Python are typically required. Excellent communication skills, critical thinking, and the ability to give constructive feedback are crucial soft skills for this role. These abilities ensure that AI systems are accurately assessed, improved, and aligned with user needs and ethical standards.

What is the difference between Ai Evaluation vs Data Analyst?

AspectAi EvaluationData Analyst
Required CredentialsTypically requires knowledge of AI/ML concepts, certifications in AI tools, and data science fundamentalsRequires degrees in statistics, data analysis, or related fields; often certifications in data analysis tools
Work EnvironmentPrimarily in tech companies, AI development teams, or research labsIn various industries including finance, marketing, healthcare, often in office settings
Employer & Industry UsageUsed by AI development firms, tech companies, and research institutionsUsed across multiple industries for data-driven decision making
Search & Comparison IntentPeople compare to understand AI-specific evaluation rolesOften compared to analyze data trends and insights

While both roles involve working with data, Ai Evaluation focuses on assessing AI models and algorithms, whereas Data Analysts interpret data to inform business decisions. Understanding these differences helps professionals choose the right career path or role based on their skills and industry needs.

What is AI evaluation?

AI evaluation refers to the process of assessing the performance, accuracy, fairness, and reliability of artificial intelligence models or systems. This involves testing AI algorithms using various metrics and datasets to ensure they meet desired standards and function as intended. Evaluators may look for issues like bias, errors, or unintended consequences. The goal is to ensure that AI systems are safe, effective, and trustworthy before they are deployed in real-world applications.
More about Ai Evaluation jobs
What cities are hiring for Ai Evaluation jobs? Cities with the most Ai Evaluation job openings:
What states have the most Ai Evaluation jobs? States with the most job openings for Ai Evaluation jobs include:
Infographic showing various Ai Evaluation job openings in the United States as of July 2026, with employment types broken down into 75% Full Time, 22% Part Time, and 3% Contract. Highlights an 66% Physical, 3% Hybrid, and 31% Remote job distribution, with an average salary of $37,137 per year, or $17.9 per hour.
AI Evaluation Scientist

$105K - $145K/yr

Other

Re-posted 4 days ago


Job description

Overview

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission requirements. This role is essential for establishing trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment. 

Contributions
  • Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics. 
  • Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time. 
  • Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs. 
  • Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations. 
  • Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements. 
  • Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports. 
  • Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety. 
  • Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments. 
  • Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability. 
  • Document evaluation processes, criteria, and results for both technical and non-technical audiences. 
  • You will contribute to the growth of our AI & Data Exploitation Practice! 
Qualifications
  • Ability to hold a position of public trust with the U.S. government. 
  • Bachelor's or Master's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field. 
  • 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred). 
  • Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems. 
  • Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain 
  • Proficiency in AI evaluation frameworks such as Ragas 
  • Experience analyzing structured and unstructured data, including text, documents, and embeddings. 
  • Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures. 
  • Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability). 
  • Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements. 
  • Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders. 
  • Experience working in agile or iterative development environments is a plus. 
  • Familiarity with OWASP LLM Top 10 Risks 
  • Relevant certifications (helpful but not required): 
    • NIST AI RMF (AISIC)
    • INFORMS CAP
    • AWS/Azure/Google ML Certifications. 
About steampunk

Steampunk relies on several factors to determine salary, including but not limited to geographic location, contractual requirements, education, knowledge, skills, competencies, and experience. The projected compensation range for this position is $105,000 to $145,000.  The estimate displayed represents a typical annual salary range for this position. Annual salary is just one aspect of Steampunk's total compensation package for employees. Learn more about additional Steampunk benefits here. 

Identity Statement

As part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.

Steampunk is a Change Agent in the Federal contracting industry, bringing new thinking to clients in the Homeland, Federal Civilian, Health and DoD sectors.  Through our Human-Centered delivery methodology, we are fundamentally changing the expectations our Federal clients have for true shared accountability in solving their toughest mission challenges.  As an employee owned company, we focus on investing in our employees to enable them to do the greatest work of their careers - and rewarding them for outstanding contributions to our growth. If you want to learn more about our story, visit http://www.steampunk.com.

Employment Type: OTHER