1

Evaluation Scientist Jobs (NOW HIRING)

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

We are looking for an AI Evaluation Scientist to design and execute evaluation processes that ensure our predictive and generative AI systems are accurate, reliable, safe, and aligned with mission ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

AI Evaluation Scientist

Mclean, VA ยท On-site

$105K - $145K/yr

The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...

The Evaluation Scientist II represents an exciting opportunity to be a member of our System Evaluation Team. The successful candidate will provide technical leadership and be responsible for projects ...

Overview The Evaluation Scientist II represents an exciting opportunity to be a member of our System Evaluation Team. The successful candidate will provide technical leadership and be responsible for ...

next page

Showing results 1-20

Evaluation Scientist information

See salary details

$37K

$91K

$160K

How much do evaluation scientist jobs pay per year?

As of Sep 12, 2026, the average yearly pay for evaluation scientist in the United States is $90,961.00, according to ZipRecruiter salary data. Most workers in this role earn between $68,000.00 and $100,000.00 per year, depending on experience, location, and employer.

What is the difference between Evaluation Scientist vs Data Analyst?

AspectEvaluation ScientistData Analyst
Required CredentialsMaster's or PhD in statistics, epidemiology, or related fieldsBachelor's or Master's in data science, statistics, or related fields
Work EnvironmentResearch settings, healthcare, government agencies, or NGOsBusiness, finance, marketing, or healthcare industries
Employer & Industry UsageUsed in research, public health, and policy evaluationUsed in business analytics, market research, and operational analysis

Evaluation Scientists focus on designing and conducting complex evaluations, often in research or public health contexts, requiring advanced degrees. Data Analysts interpret data to inform business decisions, typically with a bachelor's or master's degree. While both roles analyze data, Evaluation Scientists emphasize research design and evaluation methodologies, whereas Data Analysts focus on data processing and reporting.

What cities are hiring for Evaluation Scientist jobs?

Cities with the most Evaluation Scientist job openings:

What are popular job titles related to Evaluation Scientist jobs?

For Evaluation Scientist jobs, the most frequently searched job titles are:

AI Evaluation Scientist

Mclean, VA โ€ข On-site

Steampunk
IT Servicesย โ€ขย 201 - 500 employees

$105K - $145K/yr

Full-time

Re-posted 29 days ago


Job description

Overview

We are looking for an AI Evaluation Scientistย to design and execute evaluation processes that ensure our predictive and generative AI systems areย accurate, reliable, safe, and aligned with mission requirements. This role is essential forย establishingย trust in AI solutions and supporting continuous improvement across the AI lifecycle. The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and support responsible deployment.

Contributions
  • Implement evaluation frameworks for AI models, including accuracy, robustness, relevance, bias, hallucination rate, and safety metrics.
  • Build and maintain automated evaluation scripts, tests, and pipelines that assess AI model outputs and detect performance drift over time.ย 
  • Develop benchmark datasets, challenge sets, and scenario-based test cases tailored to mission and user needs.ย 
  • Perform structured error analysis and behavioral audits of LLMs, retrieval-augmented generation (RAG) systems, and predictive models, documenting findings and improvement recommendations.ย 
  • Collaborate with AI Developers, LLMOps Engineers, and Data Scientists to support iterative experimentation, model hardening, and quality improvements.ย 
  • Contribute to the design of human-in-the-loop evaluation workflows, integrating qualitative and quantitative insight into evaluation reports.ย 
  • Assist in mapping evaluation outcomes to responsible AI principles such as fairness, transparency, reliability, and safety.ย 
  • Partner with AI Governance Analysts to ensure evaluation outputs support compliance, documentation, and risk assessments.ย 
  • Stay current with emerging evaluation tools, frameworks, metrics, and research related to LLM assessment and generative AI reliability.ย 
  • Document evaluation processes, criteria, and results for both technical and non-technical audiences.ย 
  • You will contribute to the growth of our AI & Data Exploitation Practice!ย 
Qualifications
  • Ability to hold a position of public trust with the U.S. government.ย 
  • Bachelor's or Master's degree in Computer Science, Statistics, Machine Learning, Cognitive Science, Human-Computer Interaction, Data Science, or a related field.ย 
  • 2+ years of experience evaluating machine learning models, NLP systems, or generative AI models (LLMs preferred).ย 
  • Familiarity with evaluation metrics, statistical testing, dataset creation, and experimental design for AI systems.ย 
  • Proficiency in Python and relevant libraries such as PyTorch, Hugging Face, scikit-learn, LangChain.ย 
  • Proficiency in AI evaluation frameworks such as Ragas or DeepEval.ย 
  • Proficiency in AI traceability/observability tools based on the OpenTelemetry protocol.ย 
  • Experience analyzing structured and unstructured data, including text, documents, and embeddings.ย 
  • Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures.ย 
  • Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness, transparency, reliability).ย 
  • Strong analytical skills with the ability to interpret model weaknesses, extract insights, and recommend actionable improvements.ย 
  • Excellent written and verbal communication skills, with the ability to present evaluation findings clearly to technical and non-technical stakeholders.ย 
  • Experience working in agile or iterative development environments is a plus.ย 
  • Familiarity with OWASP LLM Top 10 Risksย 
  • Relevant certifications (helpful but not required): NIST AI RMF (AISIC), INFORMS CAP, AWS/Azure/Google ML Certifications.
  • Local to Washington, DC metro area preferred.ย 
About steampunk

Steampunk relies on several factors to determine salary, including but not limited to geographic location, contractual requirements, education, knowledge, skills, competencies, and experience. The projected compensation range for this position is $105,000 to $145,000.ย  The estimate displayed represents a typical annual salary range for this position. Annual salary is just one aspect of Steampunk's total compensation package for employees. Learn more about additional Steampunk benefits here.ย 

Identity Statement

As part of the application process, you are expected to be on camera during interviews and assessments. We reserve the right to take your picture to verify your identity and prevent fraud.

Steampunk is a Change Agent in the Federal contracting industry, bringing new thinking to clients in the Homeland, Federal Civilian, Health and DoD sectors. ย Through our Human-Centered delivery methodology, we are fundamentally changing the expectations our Federal clients have for true shared accountability in solving their toughest mission challenges.ย  If you want to learn more about our story, visit http://www.steampunk.com.

We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, or any other characteristic protected by law. Steampunk participates in the E-Verify program.ย 

Employment Type: OTHER