Applied Data Scientist, LLM Evaluation Introduction At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily ...
Applied Data Scientist, LLM Evaluation Introduction At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily ...
Applied Data Scientist, LLM Evaluation
Austin, TX ยท On-site +1
$175K - $275K/yr
Applied Data Scientist, LLM Evaluation Introduction At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily ...
Applied Data Scientist, LLM Evaluation
Austin, TX ยท On-site +1
$175K - $275K/yr
Applied Data Scientist, LLM Evaluation Introduction At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily ...
Design and implement automated systems and pipelines for evaluating LLM outputs. * Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM-based evaluations
Design and implement automated systems and pipelines for evaluating LLM outputs. * Develop metrics and KPIs to measure output quality, accuracy, and consistency using LLM-based evaluations
About Job Research Scientist, LLM Evaluation & Post-Training Company: Centific Location: Palo Alto, CA or Seattle, WA (Hybrid/Remote) Type: Full-time Role Overview As a Research Scientist, LLM ...
About Job Research Scientist, LLM Evaluation & Post-Training Company: Centific Location: Palo Alto, CA or Seattle, WA (Hybrid/Remote) Type: Full-time Role Overview As a Research Scientist, LLM ...
As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model ...
As an Applied Research Scientist, LLM Evaluation & Post-Training, you will lead research and experimentation on how evaluation design, measurement strategies, and feedback signals influence model ...
$6 - $65/hr
Our partner is looking for a AI QA Trainer - LLM Evaluation based in Netherlands. This role offers the opportunity to contribute directly to the evolution of advanced AI systems by improving their ...
New
$6 - $65/hr
Our partner is looking for a AI QA Trainer - LLM Evaluation based in Netherlands. This role offers the opportunity to contribute directly to the evolution of advanced AI systems by improving their ...
New
Senior Research Scientist, Model Evaluation
$100K - $128K/yr
... in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency. โข Build scalable and reusable tools for digging into ...
Senior Research Scientist, Model Evaluation
$100K - $128K/yr
... in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency. โข Build scalable and reusable tools for digging into ...
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
Test Engineer-AI/LLM
Palo Alto, CA ยท On-site
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
Test Engineer-AI/LLM
Palo Alto, CA ยท On-site
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
Test Engineer-AI/LLM
Palo Alto, CA ยท On-site
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
Quick apply
Test Engineer-AI/LLM
Palo Alto, CA ยท On-site
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies ...
LLM Platform Engineer
San Francisco, CA ยท On-site
$245K - $345K/yr
Create robust and scalable LLM evaluation frameworks to measure model performance, guide iteration, and prevent regression via CI/CD. * Deploy RAG systems and MCP servers to more effectively ground ...
LLM Platform Engineer
San Francisco, CA ยท On-site
$245K - $345K/yr
Create robust and scalable LLM evaluation frameworks to measure model performance, guide iteration, and prevent regression via CI/CD. * Deploy RAG systems and MCP servers to more effectively ground ...
AIML - Sr Machine Learning Engineer, Evaluation
Cupertino, CA ยท On-site
$128K - $177K/yr
We develop LLM-as-judge evaluators, train reward models calibrated against human feedback, optimize prompts and context for agents, and contribute targeted datasets and reward signals to foundation ...
AIML - Sr Machine Learning Engineer, Evaluation
Cupertino, CA ยท On-site
$128K - $177K/yr
We develop LLM-as-judge evaluators, train reward models calibrated against human feedback, optimize prompts and context for agents, and contribute targeted datasets and reward signals to foundation ...
Senior Machine Learning Engineer , LLM Evaluations
Menlo Park, CA ยท On-site
$144K - $190K/yr
About the Role Evaluation is the bottleneck in healthcare AI - you can't ship what you can't ... Design and build evaluation frameworks for LLM safety, clinical accuracy, and conversational ...
Senior Machine Learning Engineer , LLM Evaluations
Menlo Park, CA ยท On-site
$144K - $190K/yr
About the Role Evaluation is the bottleneck in healthcare AI - you can't ship what you can't ... Design and build evaluation frameworks for LLM safety, clinical accuracy, and conversational ...
LLM Applications Engineer
New York, NY ยท On-site
$130K - $175K/yr
Experience with LLM Evaluation: Knowledge of how to measure and mitigate "hallucinations" in a scientific/technical context. * Familiarity with SQL: Specifically optimizing queries that serve as the ...
LLM Applications Engineer
New York, NY ยท On-site
$130K - $175K/yr
Experience with LLM Evaluation: Knowledge of how to measure and mitigate "hallucinations" in a scientific/technical context. * Familiarity with SQL: Specifically optimizing queries that serve as the ...
AIML - Sr Machine Learning Engineer, Evaluation
$216K - $394K/yr
We develop LLM-as-judge evaluators, train reward models calibrated against human feedback, optimize prompts and context for agents, and contribute targeted datasets and reward signals to foundation ...
AIML - Sr Machine Learning Engineer, Evaluation
$216K - $394K/yr
We develop LLM-as-judge evaluators, train reward models calibrated against human feedback, optimize prompts and context for agents, and contribute targeted datasets and reward signals to foundation ...
ML/LLM Engineer - Applied AI
Austin, TX ยท On-site
Stay ahead of the curve in LLM evaluation, tuning, and agent-based architectures Qualifications * 3-6 years of experience in applied ML, with at least 1-2 years working with LLMs * Strong Python ...
ML/LLM Engineer - Applied AI
Austin, TX ยท On-site
Stay ahead of the curve in LLM evaluation, tuning, and agent-based architectures Qualifications * 3-6 years of experience in applied ML, with at least 1-2 years working with LLMs * Strong Python ...
LLM Solutions Architect
$85K - $120K/yr
Mentor engineers on LLM integration patterns, agent evaluation, and production deployment practices -- building the team's capability to own what you design. Qualifications & Skills Required * 5+ ...
Quick apply
LLM Solutions Architect
$85K - $120K/yr
Mentor engineers on LLM integration patterns, agent evaluation, and production deployment practices -- building the team's capability to own what you design. Qualifications & Skills Required * 5+ ...
AI Evaluation and Benchmarking Engineer - Remote / Telecommute
San Jose, CA ยท Remote
$55 - $60/hr
Experience benchmarking LLM agents, RL policies, autonomous agents, or hybrid AI systems. * Experience with experiment tracking, run comparison tools, metrics dashboards, or evaluation pipelines.
Quick apply
AI Evaluation and Benchmarking Engineer - Remote / Telecommute
San Jose, CA ยท Remote
$55 - $60/hr
Experience benchmarking LLM agents, RL policies, autonomous agents, or hybrid AI systems. * Experience with experiment tracking, run comparison tools, metrics dashboards, or evaluation pipelines.
Staff Machine Learning Engineer - VLM/LLM Evaluation
Mountain View, CA ยท On-site
$238K - $302K/yr
Lead the development of end-to-end evaluation systems and benchmarks for Waymo Foundation models, encompassing the entire lifecycle from pretraining and supervised fine-tuning (SFT) to reinforcement ...
Staff Machine Learning Engineer - VLM/LLM Evaluation
Mountain View, CA ยท On-site
$238K - $302K/yr
Lead the development of end-to-end evaluation systems and benchmarks for Waymo Foundation models, encompassing the entire lifecycle from pretraining and supervised fine-tuning (SFT) to reinforcement ...
Master's degree or PhD in Engineering, Computer Science, or a related technical field. * 5 years of experience with LLM evaluation, machine learning algorithms and tools, and general generative AI ...
New
Master's degree or PhD in Engineering, Computer Science, or a related technical field. * 5 years of experience with LLM evaluation, machine learning algorithms and tools, and general generative AI ...
New
Llm Evaluation information
What is the difference between Llm Evaluation vs Data Scientist?
| Aspect | Llm Evaluation | Data Scientist |
|---|---|---|
| Required Credentials | Typically requires knowledge of machine learning, NLP, and AI concepts; often a degree in computer science or related fields | Requires degrees in computer science, statistics, or related fields; often includes certifications in data analysis or machine learning |
| Work Environment | Primarily research and testing environments, focusing on AI model assessment | Data analysis, modeling, and visualization in various industries like finance, healthcare, or tech |
| Employer & Industry Usage | Used by AI research labs, tech companies, and organizations developing NLP models | Used across industries for data analysis, predictive modeling, and business insights |
While both roles involve working with data and machine learning, Llm Evaluation focuses on assessing large language models' performance, whereas Data Scientists develop and implement data-driven solutions across various sectors.

Other
Medical, Dental, Vision, Life, Retirement
Posted 3 days ago
Job description
At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily asynchronous/distributed backend server, and a frontend web application that provides a rich user experience.
About DriverWe're an early-stage startup backed by Y Combinator and Google Ventures that combines first principles technical approaches and applied LLM expertise to tackle context engineering at scale. Driver builds the context layer for employees and AI agents alike to use in developing software.
Working at DriverDriver is an early-stage but fast-growing startup. As such, we take advantage of that which startups can excel: delivery speed, flexibility, and enjoying working with a small close-knit team.
Organizational and engineering values at Driver include first-principles thinking, correct by construction, writing things down, experimentation and iteration, pragmatism, commitment to effective communication and transparency, autonomy, and ambition.
Job OverviewTitle: Applied Data Scientist, LLM Evaluation
Location: Remote or Austin, Tx
Our value is directly tied to the quality of our content at scale. The platform generates technical documentation across a complex, multi-stage pipeline - producing multiple content types at different levels of abstraction, from individual code elements up to high-level summaries. Today, changes to models, context strategies, or pipeline architecture are evaluated largely through manual review and intuition. There is no systematic way to answer: "Did this change make our output better, worse, or the same - and for which languages, repo sizes, and content types?"
This is a hard problem. LLM outputs are non-deterministic - identical inputs produce different outputs across runs, and small variations at early pipeline stages compound into meaningfully different end-user content downstream. Evaluating quality requires methodology that accounts for this: statistical reasoning over multiple runs, understanding of cascade effects through the pipeline, and rubrics that balance human judgment with automated signals.
This role builds the evaluation function from scratch. You'll define what "good" means for our generated content, build the infrastructure to measure it, and create the experimental framework that lets the team ship changes with confidence.
What You'll DoYou'll own the LLM evaluation strategy at Driver - from first principles to production infrastructure. This is a foundational role: you're not joining an existing eval team, you're building it. As the function matures, you'll seed and grow a team around it.
Define quality metrics and build evaluation datasets. Establish what "good" looks like for each content type across the pipeline. Build and curate gold-standard evaluation datasets across languages and repo archetypes (monorepos, microservices, libraries, applications). Design rubrics that capture accuracy, completeness, usefulness, and readability.
Build benchmarking and experimentation infrastructure. Create automated evaluation pipelines that score output against reference datasets. Instrument the content generation pipeline to support A/B comparisons - run the same codebase through two strategies and compare results. Build tooling for LLM-as-judge evaluation and regression detection. Integrate evaluation into CI so pipeline changes come with quality evidence.
Develop automated quality signals at scale. Build quality checks that flag degraded output without requiring human review of every document. Monitor content quality trends over time. Design sampling strategies for human review that maximize signal with minimal annotation effort.
Quantify tradeoffs and inform decisions. Run experiments on model selection, context strategies, and pipeline architecture changes. Quantify cost/quality/latency tradeoffs. Partner with the engineering team to turn evaluation insights into shipped improvements.
QualificationsEducation: Bachelor's, Master's, or PhD in Statistics, Machine Learning, Data Science, Computational Linguistics, or a related quantitative field.
Experience: Minimum 3 - 5 years in applied science, ML engineering, or data science roles with a focus on evaluation, NLP, or generative AI. 7+ years experience preferred.
Required Technical Skills
- Strong statistical foundations: experimental design, hypothesis testing, confidence intervals, effect sizes, power analysis.
- Experience designing and running evaluations for LLM or NLP systems - you've thought carefully about what "better" means when outputs are open-ended text.
- Proficient in Python and the scientific/data stack (pandas, NumPy, scipy, sklearn).
- Comfortable working in Jupyter notebooks for exploration and prototyping, and turning that work into automated pipelines.
- Experience with LLM-as-judge approaches, inter-annotator agreement, and rubric design for subjective quality assessment.
- Familiarity with the practical challenges of non-deterministic systems: variance decomposition, multi-run methodology, distinguishing signal from noise at scale.
- Strong data storytelling - you can turn experiment results into clear recommendations that drive engineering and product decisions.
Preferred and Nice-to-Have Technical Skills
- Experience with LLM APIs and prompt engineering across multiple providers.
- Familiarity with evaluation frameworks (e.g., RAGAS, DeepEval, custom harnesses).
- Experience building data pipelines or ETL workflows (Airflow, Dagster, or similar).
- Comfort with SQL and working directly against production data stores.
- Experience with visualization tools (Matplotlib, Plotly, Streamlit) for building internal dashboards and reports.
- Background in code understanding, developer tools, or technical documentation.
- Experience building or managing annotation pipelines and human evaluation workflows.
- Competitive Compensation Packages - Cash & Equity
- Flexible Work Culture
- Unlimited Time Off + 12 Paid Company Holidays
- Insurance - Health, Dental, & Vision
- Life Insurance & FSA Accounts
- 401(k) Retirement Accounts - Traditional, Roth, or Both
- Quarterly Team Offsites
Driver is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.