Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Experience in applied AI research, machine learning, model evaluation, or data-centric AI ... Strong understanding of benchmark design, human evaluation, rubric development, and statistical ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Advising R&D partnerships developing artificial intelligence products/tools to improve model function, design clinical validation studies, and implement production deployment and maintenance ...
Advising R&D partnerships developing artificial intelligence products/tools to improve model function, design clinical validation studies, and implement production deployment and maintenance ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Advising R&D partnerships developing artificial intelligence products/tools to improve model function, design clinical validation studies, and implement production deployment and maintenance ...
Quick apply
Advising R&D partnerships developing artificial intelligence products/tools to improve model function, design clinical validation studies, and implement production deployment and maintenance ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment. ■ About Kotoba Kotoba is a generative AI company on a ...
Integrate your research with platform, firmware and software teams to ensure your work reaches the robot rather than remaining purely experimental or theoretical. * Contribute to sim-to-real transfer ...
Integrate your research with platform, firmware and software teams to ensure your work reaches the robot rather than remaining purely experimental or theoretical. * Contribute to sim-to-real transfer ...
... and applied doctoral research. Anderson University is particularly interested in scholar ... History and Development * LEAI 740 - Machine Learning and Generative AI * LEAI 750 - AI Ethics and ...
... and applied doctoral research. Anderson University is particularly interested in scholar ... History and Development * LEAI 740 - Machine Learning and Generative AI * LEAI 750 - AI Ethics and ...
Artificial Intelligence Research Development information
See salary details
$48.5K - $59.1K
9% of jobs
$59.1K - $69.7K
9% of jobs
$77.6K is the 25th percentile. Wages below this are outliers.
$69.7K - $80.3K
8% of jobs
$80.3K - $90.9K
11% of jobs
$90.9K - $101.5K
9% of jobs
The median wage is $103.1K / yr.
$101.5K - $112K
14% of jobs
$117.3K is the 75th percentile. Wages above this are outliers.
$112K - $122.6K
27% of jobs
$122.6K - $133.2K
5% of jobs
$133.2K - $143.8K
2% of jobs
$143.8K - $154.4K
1% of jobs
$154.4K - $165K
3% of jobs
$48.5K
$101.8K
$165K
How much do artificial intelligence research development jobs pay per year?
What cities are hiring for Artificial Intelligence Research Development jobs?
Cities with the most Artificial Intelligence Research Development job openings:
What states have the most Artificial Intelligence Research Development jobs?
States with the most job openings for Artificial Intelligence Research Development jobs include:
What are popular job titles related to Artificial Intelligence Research Development jobs?
For Artificial Intelligence Research Development jobs, the most frequently searched job titles are:

Artificial Intelligence Researcher
Fremont, CA • On-site
Other
Posted 10 days ago
Job description
Location: San Francisco
Employment Type: Full-time
Work Model: In-person
About Verita AI
Verita AI works with leading AI companies to identify model gaps and build the human data needed to improve model performance. Verita AI operates a vetted expert network that connects specialized professionals with leading AI laboratories and human-data companies. The network comprises more than 5,000 experts across finance, medicine, law, engineering, music, and other professional domains.
We recently raised a $6 million seed round led by Kindred Ventures.
About the Role
We are hiring an Applied AI Researcher to work directly with clients on model evaluation and data strategy.
You will evaluate model performance, identify failure modes, and recommend the datasets, rubrics, expert workflows, and quality controls needed to address them. You will then work with our operations and engineering teams to turn these recommendations into scalable data programs.
What You’ll Do
- Work with clients to understand their models, goals, and performance gaps.
- Design evaluations for generative, multimodal, reasoning, tool-use, and agentic AI systems.
- Analyze model outputs and benchmark results to identify and quantify failure modes.
- Recommend data solutions such as supervised fine-tuning data, preference data, expert demonstrations, critiques, and evaluation datasets.
- Write client proposals covering the methodology, data design, quality controls, staffing, deliverables, and expected impact.
- Create annotation guidelines, scoring rubrics, gold-standard tasks, and evaluator-training programs.
- Design pilot studies and measure whether data interventions improve model performance.
- Build quality systems using calibration tasks, blind review, adjudication, and expert scoring.
- Work with operations and engineering teams to launch and scale data pipelines.
- Present findings and recommendations to clients.
What We’re Looking For
- Experience in applied AI research, machine learning, model evaluation, or data-centric AI.
- Experience evaluating foundation models or generative AI systems.
- Strong understanding of benchmark design, human evaluation, rubric development, and statistical analysis.
- Ability to translate model failures into practical data solutions.
- Strong Python skills and experience working with model APIs and structured datasets.
- Familiarity with supervised fine-tuning, preference optimization, RLHF/RLAIF, reward modeling, synthetic data, or LLM-as-a-judge evaluation.
- Strong technical writing and client communication skills.
- Ability to independently structure and execute ambiguous research projects.
Nice to Have
- Experience at an AI lab, foundation-model company, AI data company, or post-training team.
- Experience designing expert-data or human-evaluation programs.
- Experience evaluating multimodal, coding, agentic, or tool-use systems.
- Publications at conferences such as NeurIPS, ICML, ICLR, ACL, or EMNLP.
- Previous client-facing research, consulting, solutions engineering, or forward-deployed experience.
- Public research, code, benchmarks, or evaluation frameworks.
Please make sure you have any relevant work samples, including model evaluations, benchmarks, error analyses, research, technical writing, or code repositories.