You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions. * You will create, configure, and monitor data annotation jobs to keep evaluation and ...
You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions. * You will create, configure, and monitor data annotation jobs to keep evaluation and ...
Develop LLM-based knowledge graph construction pipelines that extract and link citations, entities ... Annotation workflow design and evaluation framework development for document understanding tasks ...
Develop LLM-based knowledge graph construction pipelines that extract and link citations, entities ... Annotation workflow design and evaluation framework development for document understanding tasks ...
AI Evaluation Engineer
Kitchener, ON · On-site
You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions. * You will create, configure, and monitor data annotation jobs to keep evaluation and ...
Quick apply
AI Evaluation Engineer
Kitchener, ON · On-site
You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions. * You will create, configure, and monitor data annotation jobs to keep evaluation and ...
Llm Annotation information
What is LLM annotation?
What are the key skills and qualifications needed to thrive as an LLM annotation specialist?
What are some common challenges faced by LLM annotation specialists, and how can they be addressed?
What is the difference between Llm Annotation vs Data Labeler?
| Aspect | Llm Annotation | Data Labeler |
|---|---|---|
| Required Credentials | Basic computer skills, sometimes familiarity with AI tools | Basic skills, often on-the-job training |
| Work Environment | Remote or office-based, tech-focused | Remote or on-site, varied industries |
| Industry Usage | AI, machine learning, NLP projects | Various industries including marketing, healthcare, and tech |
| Search & Comparison Intent | Understanding roles in AI data preparation | General data labeling tasks |
In summary, Llm Annotation involves specialized annotation for large language models, often requiring familiarity with AI tools, while Data Labeler is a broader role focused on labeling data across multiple industries with minimal technical requirements.
How to become an Llm annotator?
What are popular job titles related to Llm Annotation jobs in Ontario?
For Llm Annotation jobs in Ontario, the most frequently searched job titles are:
What job categories do people searching Llm Annotation jobs in Ontario look for?
The top searched job categories for Llm Annotation jobs in Ontario are:

Job description
Your role
As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside our existing evaluation lead. A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions.
This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office.
What you'll do
- You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
- You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward.
- You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
- You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule.
- You will develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams.
- You will investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis.
- You will collaborate with cross-functional teams, including applied science, engineering, and Product QA.
Skills you'll bring
- Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field.
- 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products.
- Experience designing structured test strategies across manual and automated workflows.
- Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products.
- Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
- Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns.
- Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.
About Dialpad
Sourced by ZipRecruiter
Industry
Telecommunications
Company size
1,001 - 5,000 Employees
Headquarters location
San Francisco, CA, US
Year founded
2011