As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model ...
As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model ...
Applied Data Scientist, LLM Evaluation
Austin, TX · On-site +1
$175K - $275K/yr
Applied Data Scientist, LLM Evaluation Introduction At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily ...
Applied Data Scientist, LLM Evaluation
Austin, TX · On-site +1
$175K - $275K/yr
Applied Data Scientist, LLM Evaluation Introduction At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily ...
As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model ...
As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model ...
Applied Data Scientist, LLM Evaluation Introduction At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily ...
Applied Data Scientist, LLM Evaluation Introduction At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily ...
Design and implement automated systems and pipelines for evaluating LLM outputs. * Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM-based evaluations
Design and implement automated systems and pipelines for evaluating LLM outputs. * Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM-based evaluations
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
Test Engineer-AI/LLM
Palo Alto, CA · On-site
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
Quick apply
Test Engineer-AI/LLM
Palo Alto, CA · On-site
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
Test Engineer-AI/LLM
Palo Alto, CA · On-site
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
Test Engineer-AI/LLM
Palo Alto, CA · On-site
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
LLM Specialist
$16.50 - $21.75/hr
This role is responsible for evaluating, selecting, implementing, securing, and optimizing LLM solutions that support business objectives, enhancing member experiences, improving operational ...
LLM Specialist
$16.50 - $21.75/hr
This role is responsible for evaluating, selecting, implementing, securing, and optimizing LLM solutions that support business objectives, enhancing member experiences, improving operational ...
Senior Machine Learning Engineer , LLM Evaluations
Menlo Park, CA · On-site
$144K - $190K/yr
Design and build evaluation frameworks for LLM safety, clinical accuracy, and conversational quality * Develop synthetic data generation pipelines to stress-test models across diverse clinical ...
Senior Machine Learning Engineer , LLM Evaluations
Menlo Park, CA · On-site
$144K - $190K/yr
Design and build evaluation frameworks for LLM safety, clinical accuracy, and conversational quality * Develop synthetic data generation pipelines to stress-test models across diverse clinical ...
LLM Platform Engineer
San Francisco, CA · On-site
$245K - $345K/yr
Create robust and scalable LLM evaluation frameworks to measure model performance, guide iteration, and prevent regression via CI/CD. * Deploy RAG systems and MCP servers to more effectively ground ...
LLM Platform Engineer
San Francisco, CA · On-site
$245K - $345K/yr
Create robust and scalable LLM evaluation frameworks to measure model performance, guide iteration, and prevent regression via CI/CD. * Deploy RAG systems and MCP servers to more effectively ground ...
LLM Platform Engineer
San Francisco, CA · On-site
$245K - $345K/yr
Create robust and scalable LLM evaluation frameworks to measure model performance, guide iteration, and prevent regression via CI/CD. * Deploy RAG systems and MCP servers to more effectively ground ...
LLM Platform Engineer
San Francisco, CA · On-site
$245K - $345K/yr
Create robust and scalable LLM evaluation frameworks to measure model performance, guide iteration, and prevent regression via CI/CD. * Deploy RAG systems and MCP servers to more effectively ground ...
LLM Optimization & Evaluation
Austin, TX · On-site
Info Way Solutions is seeking a highly skilled LLM Optimization & Evaluation Specialist to design, optimize, and evaluate large language model systems in production. This role focuses on prompt ...
LLM Optimization & Evaluation
Austin, TX · On-site
Info Way Solutions is seeking a highly skilled LLM Optimization & Evaluation Specialist to design, optimize, and evaluate large language model systems in production. This role focuses on prompt ...
LLM Applications Engineer
New York, NY · On-site
$130K - $175K/yr
Experience with LLM Evaluation: Knowledge of how to measure and mitigate "hallucinations" in a scientific/technical context. * Familiarity with SQL: Specifically optimizing queries that serve as the ...
LLM Applications Engineer
New York, NY · On-site
$130K - $175K/yr
Experience with LLM Evaluation: Knowledge of how to measure and mitigate "hallucinations" in a scientific/technical context. * Familiarity with SQL: Specifically optimizing queries that serve as the ...
ML/LLM Engineer - Applied AI
Austin, TX · On-site
... LLM evaluation, tuning, and agent-based architectures. Qualifications : Required : • 3-6 years of experience in applied ML, with at least 1-2 years working with LLMs • Strong Python skills and ...
ML/LLM Engineer - Applied AI
Austin, TX · On-site
... LLM evaluation, tuning, and agent-based architectures. Qualifications : Required : • 3-6 years of experience in applied ML, with at least 1-2 years working with LLMs • Strong Python skills and ...
LLM Solutions Architect
$85K - $120K/yr
Mentor engineers on LLM integration patterns, agent evaluation, and production deployment practices - building the team's capability to own what you design. Qualifications & Skills Required * 5+ ...
LLM Solutions Architect
$85K - $120K/yr
Mentor engineers on LLM integration patterns, agent evaluation, and production deployment practices - building the team's capability to own what you design. Qualifications & Skills Required * 5+ ...
LLM Solutions Architect
Concord, NC · On-site
$85K - $120K/yr
Mentor engineers on LLM integration patterns, agent evaluation, and production deployment practices -- building the team's capability to own what you design. Qualifications & Skills Required * 5+ ...
Quick apply
LLM Solutions Architect
Concord, NC · On-site
$85K - $120K/yr
Mentor engineers on LLM integration patterns, agent evaluation, and production deployment practices -- building the team's capability to own what you design. Qualifications & Skills Required * 5+ ...
Senior Software Engineer - LLM Evaluation
San Francisco, CA · On-site
$110 - $165/hr
Senior Software Engineer - LLM Evaluation Skills JavaScript JavaScript Python Overview About Turing Based in San Francisco, California, Turing is the world's leading research accelerator for frontier ...
Senior Software Engineer - LLM Evaluation
San Francisco, CA · On-site
$110 - $165/hr
Senior Software Engineer - LLM Evaluation Skills JavaScript JavaScript Python Overview About Turing Based in San Francisco, California, Turing is the world's leading research accelerator for frontier ...
ML/LLM Engineer - Applied AI
Austin, TX · On-site
Stay ahead of the curve in LLM evaluation, tuning, and agent-based architectures Qualifications * 3-6 years of experience in applied ML, with at least 1-2 years working with LLMs * Strong Python ...
ML/LLM Engineer - Applied AI
Austin, TX · On-site
Stay ahead of the curve in LLM evaluation, tuning, and agent-based architectures Qualifications * 3-6 years of experience in applied ML, with at least 1-2 years working with LLMs * Strong Python ...
Senior Research Scientist, Model Evaluation
$100K - $128K/yr
... in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency. • Build scalable and reusable tools for digging into ...
Senior Research Scientist, Model Evaluation
$100K - $128K/yr
... in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency. • Build scalable and reusable tools for digging into ...
Llm Evaluator information
See salary details
$29.5K - $36.5K
13% of jobs
$36.5K - $43.5K
12% of jobs
$43.6K is the 25th percentile. Wages below this are outliers.
$43.5K - $50.5K
14% of jobs
The median wage is $56.6K / yr.
$50.5K - $57.5K
13% of jobs
$57.5K - $64.5K
12% of jobs
$64.5K - $71.5K
6% of jobs
$75.7K is the 75th percentile. Wages above this are outliers.
$71.5K - $78.5K
9% of jobs
$78.5K - $85.5K
11% of jobs
$85.5K - $92.5K
4% of jobs
$92.5K - $99.5K
4% of jobs
$99.5K - $106.5K
2% of jobs
$29.5K
$65.5K
$106.5K
How much do llm evaluator jobs pay per year?
What is the difference between Llm Evaluator vs Data Annotator?
| Aspect | Llm Evaluator | Data Annotator |
|---|---|---|
| Required Credentials | Typically requires knowledge of AI, NLP, or machine learning; often a degree in computer science or related field | Usually requires attention to detail; high school diploma or equivalent often sufficient |
| Work Environment | Primarily office or remote work focused on evaluating AI model outputs | Often in a data labeling or annotation environment, sometimes remote |
| Employer & Industry Usage | Used in AI and tech companies to assess language model performance | Used across industries for preparing training data for machine learning models |
While both roles involve working with data and AI, Llm Evaluators focus on assessing and improving language models' outputs, requiring technical knowledge. Data Annotators primarily label data to train models, often with less technical background. Understanding these differences helps clarify career paths and employer expectations in AI development.
What cities are hiring for Llm Evaluator jobs?
Cities with the most Llm Evaluator job openings:
What states have the most Llm Evaluator jobs?
States with the most job openings for Llm Evaluator jobs include:
What job categories do people searching Llm Evaluator jobs look for?
The top searched job categories for Llm Evaluator jobs are:

Full-time
Re-posted 24 days ago
Innodata rating
7.5
Based on 6 frontline employees who took The Breakroom Quiz
162nd of 244 rated software companies
Job description
Scope of the Role:
Innodata is expanding its GenAI research capability to advance state-of-the-art evaluation and post-training methods for LLM and multimodal systems. As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model improvement.
This role is ideal for a technically rigorous researcher who is deeply fluent in modern LLM evaluation and post-training, and who can turn research insight into practical methods for customer solutions and internal platform innovation. You will work across human-in-the-loop and AI-augmented workflows, partnering with Language Data Scientists and AI/ML Research Engineers to design and validate evaluation frameworks that drive measurable model gains.
The ideal candidate combines strong experimental and statistical judgment with hands-on technical ability and can engage as a peer with research and engineering stakeholders at leading AI companies.
What You'll Own:
As an Applied Research Scientist, LLM Evaluation & Post-Training, you will help define the next generation of evaluation-driven model improvement workflows. You will study how different evaluation approaches (human, automated, hybrid) shape model selection and post-training outcomes, and you will design experiments that produce credible, actionable conclusions.
Your work may include designing benchmark datasets, developing evaluation taxonomies and protocols, defining metrics and scoring methodologies, analyzing failure modes, and testing how changes in evaluation setup affect downstream fine-tuning results. You will also support customer engagements by bringing scientific rigor to evaluation strategy, methodology review, and technical recommendations.
This is a highly collaborative role that sits at the intersection of research, engineering, and language/data operations. Additional responsibilities include (but are not limited to):
- Define and execute a research agenda focused on LLM evaluation and post-training, especially evaluation-driven model improvement
- Design rigorous experiments to study how evaluation methodologies impact fine-tuning and post-training outcomes
- Develop and validate evaluation frameworks for LLM and multimodal systems, including:
- benchmark/task design
- scoring methods
- judge/model-assisted evaluation
- human evaluation protocols
- robustness/stress testing
- Lead research on advanced evaluation domains, including long-context, cross-modal, and dynamic multi-turn evaluations
- Study the effectiveness and limitations of existing evaluation techniques, and propose improved methodologies with clear validity and scalability tradeoffs
- Analyze model behavior and failure patterns; generate actionable recommendations for model improvement and evaluation redesign
- Collaborate with AI/ML Research Engineers to translate research methods into scalable evaluation and post-training pipelines
- Collaborate with Language Data Scientists to integrate human-in-the-loop and synthetic data/evaluation strategies into research programs
- Engage with customer technical stakeholders to understand evaluation goals, review methodologies, and provide expert recommendations
- Contribute to internal benchmark datasets, evaluation frameworks, and reusable research assets
- Produce high-quality technical documentation, internal research reports, and client-facing materials explaining methods, results, assumptions, and limitations
- Contribute to thought leadership and best practices in LLM evaluation, post-training, and GenAI quality measurement
You'll Thrive in This Role If You Have:
- MS/PhD in Computer Science, Machine Learning, Statistics, Applied Mathematics, AI, or a related quantitative scientific field (PhD strongly preferred)
- 5+ years of relevant experience in applied research / research science in ML/AI, with substantial work in LLMs or foundation models
- Demonstrated experience with LLM evaluation, benchmarking, alignment, post-training, or model quality research
- Strong foundation in experimental design, statistical analysis, and scientific reasoning for ML systems
- Strong coding skills in Python for research experimentation and analysis (e.g., data processing, evaluation pipelines, statistical analysis, visualization)
- Experience working with modern ML tooling/frameworks (e.g., PyTorch, Hugging Face, JAX/TensorFlow as applicable) sufficient to design and execute model/evaluation experiments
- Ability to evaluate and compare human and automated evaluation methods, including tradeoffs in cost, reliability, validity, and scalability
- Experience designing evaluation studies and protocols that are reproducible across datasets, model versions, and evaluation runs
- Ability to collaborate directly with technical stakeholders including research scientists, ML engineers, data scientists, and customer technical counterparts
- Strong communication skills and ability to present nuanced technical conclusions, assumptions, and limitations clearly
The expected salary range for this position is $175,000 - $225,000 USD per year, based on experience, skills, and qualifications.
About Innodata
Sourced by ZipRecruiter
Industry
It services
Company size
501 - 1,000 Employees
Headquarters location
Hackensack, NJ, US
Year founded
1988