1

Evaluation Research Jobs (NOW HIRING)

Work with Population Research, Evaluation Research, and Simulation Engineering to turn validated methods into reliable simulation and product capabilities. * Communicate findings plainly, including ...

New

Opportunity to lead organization-wide research and evaluation initiatives that directly impact visitor experiences, educational programs, and organizational strategy * Collaborative, mission-driven ...

$7.0K/mo

Participates in program evaluation research activities; Reports allegations of child abuse, neglect and/or exploitation to authorities; Provides training to Department staff and interns, as well as ...

Showing results 41-60

Evaluation Research information

See salary details

$37K

$106K

$142.5K

How much do evaluation research jobs pay per year?

As of Aug 15, 2026, the average yearly pay for evaluation research in the United States is $106,012.00, according to ZipRecruiter salary data. Most workers in this role earn between $104,000.00 and $104,000.00 per year, depending on experience, location, and employer.

What is evaluation research?

Evaluation research is a systematic method of assessing the design, implementation, and outcomes of programs, policies, or projects to determine their effectiveness and inform decision-making. It combines quantitative and qualitative research methods to collect and analyze data, providing evidence-based insights about what works and why. Evaluation research is commonly used in fields such as education, healthcare, social services, and public policy to improve program efficiency and impact.

What types of projects or sectors do evaluation researchers typically work on, and how does this diversity impact daily responsibilities?

Evaluation researchers often work across a variety of sectors, including education, public health, social services, and nonprofit organizations. This diversity means that your daily tasks can range from designing data collection tools and conducting field interviews to analyzing large datasets and presenting findings to stakeholders. The role requires adaptability, as each project may have different methodologies, objectives, and reporting requirements. Collaborating with subject matter experts and program managers is common, making strong communication and project management skills essential.

What are the key skills and qualifications needed to thrive in evaluation research, and why are they important?

To succeed in Evaluation Research, you need a solid background in research methods, data analysis, and program evaluation, often supported by a degree in social sciences or a related field. Proficiency with statistical software such as SPSS, R, or NVivo, and familiarity with survey tools and data collection platforms are typically required. Strong written and verbal communication, critical thinking, and attention to detail help present findings clearly and work effectively with stakeholders. These skills ensure rigorous, actionable evaluations that inform decision-making and improve program outcomes.

What is the difference between Evaluation Research vs Data Analyst?

AspectEvaluation ResearchData Analyst
Required CredentialsMaster's degree in social sciences, research methods, or related fieldsBachelor's or Master's in Data Science, Statistics, or related fields
Work EnvironmentResearch institutions, government agencies, NGOsCorporate, finance, healthcare, tech companies
Employer & Industry UsageUsed in program assessment, policy analysis, social researchUsed in data interpretation, reporting, business insights

Evaluation Research focuses on assessing programs and policies through systematic research, often in social or public sectors. Data Analysts interpret and visualize data to support business decisions. While both roles require analytical skills, Evaluation Research emphasizes research design and evaluation methods, whereas Data Analysts focus on data manipulation and reporting.

More about Evaluation Research jobs

What cities are hiring for Evaluation Research jobs?

Cities with the most Evaluation Research job openings:

What are the most commonly searched types of Evaluation Research jobs?

The most popular types of Evaluation Research jobs are:

What states have the most Evaluation Research jobs?

States with the most job openings for Evaluation Research jobs include:

Infographic showing various Evaluation Research job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 86% Full Time, 11% Part Time, and 2% Contract. Highlights an 87% Physical, 3% Hybrid, and 10% Remote job distribution, with an average salary of $106,012 per year, or $51 per hour.

Head of Prediction Research

Aaru

New York, NY โ€ข On-site

Full-time

Posted 2 days ago

New


Job description

About Aaru
Aaru builds simulations of human behavior. Each simulation contains a population of agents, each representing a person who could plausibly exist in the real world and capable of making decisions within a modeled environment. Companies and institutions use these simulations to test consequential choices before committing, from product launches and policy changes to critical communications. Because the agents are simulated rather than recruited, they can reason through complex hypotheticals without fatigue or the response effects common in human studies.
The role
Prediction Research builds systems that estimate future or otherwise unknown outcomes from data. The work spans forecasts of specific events, likely behavior across a population, and conditional predictions that update as the available information changes. Some systems will work directly from structured data. Others will use agents to retrieve evidence, call tools, compare hypotheses, and reason through a problem before producing an estimate.
As Head of Prediction Research, you will set the scientific direction for this work and lead it from early experiments through validated systems. You will decide which prediction problems matter, what evidence is needed to solve them, and which technical approaches merit sustained investment. Aaru's data organization will sit within this function, giving you responsibility for the data strategy and data products that support prediction across the company.
This is a hands-on research leadership role. You will write code, design experiments, inspect individual failures, and work closely with researchers and engineers on the most important technical questions. You will also recruit and lead a small team capable of making progress on problems that do not yet have established methods or benchmarks.
What you will do
  • Define a focused research agenda for forecasting, behavioral prediction, calibration, and agentic prediction systems.
  • Build and test systems that combine language models with structured data, retrieval, tools, and quantitative methods.
  • Develop prediction methods from real-world records such as transactions, product usage, event histories, operational data, market data, and longitudinal outcomes.
  • Improve estimates of population behavior and how those behaviors change with a person's attributes, prior behavior, or surrounding context.
  • Determine when language-model reasoning or explicit agent simulation adds predictive signal beyond simpler statistical and machine learning approaches.
  • Establish Aaru's data strategy and lead the team responsible for acquiring, joining, cleaning, documenting, and serving research-quality data.
  • Create feedback loops in which resolved events and customer outcomes improve future systems while clean evaluation sets remain protected.
  • Design experiments with strong baselines and temporal holdouts, then test whether gains persist across domains, time periods, and populations.
  • Work with Population Research, Evaluation Research, and Simulation Engineering to turn validated methods into reliable simulation and product capabilities.
  • Communicate findings plainly, including negative results, unstable improvements, and cases where the evidence does not support a confident prediction.
  • Hire, mentor, and lead exceptional researchers, research engineers, and data specialists while remaining a direct contributor.
Representative research directions
  • Forecast a future event or business outcome using only the information available at the time the prediction would have been made.
  • Predict demand, adoption, conversion, retention, purchasing behavior, or other outcomes from transaction and usage data, including for groups with limited direct history.
  • Build accurate generators for variables that matter to a simulation and measure how errors in those marginals affect downstream results.
  • Produce conditional predictions that respond coherently to a price change, product launch, information event, or shift in the economic environment.
  • Develop an agentic forecasting system that retrieves evidence, decomposes a question, tests assumptions, and revises its estimate before returning a calibrated probability.
  • Compare direct prediction with population-based simulation to identify when modeling individual agents produces a meaningful gain.
  • Build methods for rare, novel, or rapidly changing settings where historical labels are sparse and standard supervised learning is unreliable.
  • Study how prediction quality changes with model capability, inference-time computation, retrieval quality, data coverage, and historical context.
  • Improve calibration and selective prediction so the system can express uncertainty and decline to make claims when the evidence is weak.
How we work
We treat prediction as an empirical science. Progress is measured against future or otherwise held-out outcomes, with particular attention to calibration, behavioral shift, subgroup performance, and data leakage. Strong baselines matter, including simple historical rates and conventional statistical models.
Research at Aaru is exploratory, but it must eventually change what the company can build or what it believes. Papers and benchmarks can be useful along the way. The central goal is to produce prediction systems that remain useful when they encounter new data, new customers, and real consequences.
You might thrive in this role if
  • You have developed an original research agenda in frontier machine learning, forecasting, probabilistic modeling, agentic systems, or an environment with a comparable bar for rigor and ambition.
  • You have built predictive systems from messy, heterogeneous data and tested them against observed outcomes.
  • You are comfortable moving between research strategy, statistical reasoning, model design, data design, implementation, and detailed error analysis.
  • You understand the strengths and failure modes of language models and can combine them effectively with structured data and quantitative methods.
  • You can turn a poorly specified prediction problem into a sequence of experiments that resolves the most important uncertainties.
  • You care deeply about calibration, temporal validity, selection effects, leakage, and performance under condition shift.
  • You have led researchers or a major technical direction while continuing to contribute directly to the work.
  • You communicate results clearly and change direction when the evidence contradicts an attractive idea.
  • You want to build in person, in New York, at high speed.
Strong candidates may also have
  • Work in time-series modeling, econometrics, decision science, quantitative social science, recommender systems, risk modeling, or causal inference.
  • Experience with LLM agents, retrieval, tool use, post-training, synthetic environments, or inference-time scaling.
  • Experience building proprietary datasets, data products, acquisition programs, or learning systems that improve as outcomes resolve.
  • A record of research that improved an operational prediction system or changed a consequential product or business decision.
  • Experience recruiting and mentoring unusually strong researchers, research engineers, or data scientists.
Success in this role looks like
  • Aaru has a clear prediction research agenda organized around a small number of important, testable questions.
  • New methods outperform strong baselines on held-out and prospective outcomes, with well-understood limits.
  • Aaru's data becomes a compounding research advantage with clear provenance and direct value to model development.
  • Forecasts and other predictive outputs are calibrated, useful in real decisions, and honest about uncertainty.
  • Validated methods move into production and improve the quality of Aaru's simulations and customer-facing products.
  • A small, exceptional team develops a reputation for prediction research that is technically ambitious, empirically serious, and useful in the world.