Dialpad
Dialpad

30 Dialpad Data Analyst Jobs Hiring Near You

You will create, configure, and monitor data annotation jobs to keep evaluation and calibration ... Strong analytical skills for investigating failures, comparing outputs, and identifying actionable ...

Agentic Transformation Advisor

Austin, TX · On-site

$20.75 - $27.75/hr

... Dialpad's agentic products across new and recently closed customers. You'll work closely with ... Strong business judgment and analytical skills, with the ability to translate data into clear ...

Implement integrations with customer data sources, APIs, SaaS platforms, CRMs, ERPs, and enterprise ... Run iterative testing, failure analysis, and performance tuning to improve quality, groundedness ...

New

Own the data strategy underneath it all: acquisition, consent and usage rights, sampling, and ... Sit inside eval reviews, error analyses, and incident retros as a peer. You should be able to look ...

Partner Marketing Coordinator

Austin, TX · On-site

$42K - $58K/yr

Solid analytical skills with experience tracking KPIs and supporting data-informed improvements to partner programs. * Working knowledge of Salesforce and marketing automation tools, including ...

Showing results 21-30

AI Evaluation Engineer

Dialpad

Vancouver, BC • On-site

Full-time

Re-posted 13 days ago


Job description

Your role
As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside our existing evaluation lead. A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions.

This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office.

What you'll do 

  • You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
  • You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward.
  • You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
  • You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule.
  • You will develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams.
  • You will investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis.
  • You will collaborate with cross-functional teams, including applied science, engineering, and Product QA.

Skills you'll bring 

  • Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field.
  • 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products.
  • Experience designing structured test strategies across manual and automated workflows.
  • Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products.
  • Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
  • Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns.
  • Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.