1

Evaluation Engineer Jobs in Colorado (NOW HIRING)

Hybrid 3 days (offices in NYC, Denver, CO and Charlotte, NC area) Position Summary As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to ...

next page

Showing results 1-20

Evaluation Engineer information

See Colorado salary details

$58.4K

$135.1K

$174.6K

How much do evaluation engineer jobs pay per year?

As of Aug 11, 2026, the average yearly pay for evaluation engineer in Colorado is $135,067.00, according to ZipRecruiter salary data. Most workers in this role earn between $124,600.00 and $157,700.00 per year, depending on experience, location, and employer.

What is an evaluation engineer?

An Evaluation Engineer is a professional who assesses products, systems, or processes to ensure they meet specified standards and performance criteria. They are responsible for designing and conducting tests, analyzing results, and recommending improvements or changes. Evaluation Engineers work in various industries, including manufacturing, electronics, software, and automotive, to support product development and quality assurance. Their work helps companies deliver reliable and effective products to the market.

What are some common challenges faced by evaluation engineers when assessing new products or systems?

Evaluation Engineers often encounter challenges such as tight project deadlines, rapidly evolving technology, and the need to balance thorough testing with efficiency. They may also face difficulties in obtaining comprehensive data or replicating real-world scenarios during evaluations. Collaborating closely with cross-functional teams—like design, manufacturing, and quality assurance—is essential to address these challenges and ensure accurate, actionable results.

What is the difference between Evaluation Engineer vs Test Engineer?

AspectEvaluation EngineerTest Engineer
Required CredentialsBachelor's in Engineering, certifications in testing or evaluation methodsBachelor's in Engineering, certifications in testing or quality assurance
Work EnvironmentResearch labs, product development, quality assessmentManufacturing plants, testing labs, product validation
Industry UsageUsed in electronics, aerospace, automotive for evaluating performanceUsed across industries for testing products and systems

Evaluation Engineers focus on assessing product performance, reliability, and compliance through detailed analysis, often in research or development settings. Test Engineers primarily execute testing procedures to identify defects and ensure quality during manufacturing or pre-release stages. While both roles require technical skills and certifications, Evaluation Engineers emphasize evaluation and analysis, whereas Test Engineers concentrate on testing execution and defect detection.

What are the key skills and qualifications needed to thrive as an evaluation engineer, and why are they important?

To thrive as an Evaluation Engineer, you need a solid background in engineering principles, analytical problem-solving, and experience with product testing, often supported by a degree in engineering or a related field. Familiarity with testing equipment, data analysis tools (such as MATLAB or LabVIEW), and industry-specific standards or certifications is typically required. Strong attention to detail, effective communication, and collaboration skills help Evaluation Engineers accurately assess products and share findings with cross-functional teams. These skills are crucial for ensuring product quality, safety, and compliance with regulatory and customer requirements.
What are popular job titles related to Evaluation Engineer jobs in Colorado? For Evaluation Engineer jobs in Colorado, the most frequently searched job titles are:
What job categories do people searching Evaluation Engineer jobs in Colorado look for? The top searched job categories for Evaluation Engineer jobs in Colorado are:
What cities in Colorado are hiring for Evaluation Engineer jobs? Cities in Colorado with the most Evaluation Engineer job openings:
Infographic showing various Evaluation Engineer job openings in Colorado as of August 2026, with employment types broken down into 90% Full Time, 5% Part Time, and 5% Contract. Highlights an 86% Physical, 5% Hybrid, and 9% Remote job distribution, with an average salary of $135,067 per year, or $64.9 per hour.

AI Evaluation Engineer

Capital Rx

Denver, CO • On-site

Other

This job post has expired today. Applications are no longer accepted.


Job description

About Judi Health

Judi Health is a health technology company providing benefit administration solutions to employers, unions, health plans, and government entities. Judi Health replaces fragmented, outdated systems with the industry's first Unified Claims Processing architecture, seamlessly consolidating pharmacy and medical benefit administration on a single, secure platform. By delivering true price transparency, eliminating unnecessary middleman fees, and leveraging advanced AI-powered care delivery, Judi Health helps clients achieve unprecedented operational efficiency and service levels.
At Judi Health, we're deploying the infrastructure our country needs to deliver the healthcare we all deserve. We are the intelligence platform powering benefits plans for millions of Americans and proudly leading the next generation of care. To learn more, visit www.judi.health.

Hybrid 3 days (offices in NYC, Denver, CO and Charlotte, NC area)

Position Summary

As an AI Evaluation Engineer at Judi Health, you will build the testing frameworks, metrics, and tooling used to assess the safety, reliability, and accuracy of AI models and autonomous agents in production. This role bridges the gap between model development and realworld usage by translating ambiguous product goals into measurable quality targets.

We're looking for someone to lead evaluation end-to-end - from unit and integration testing to offline, online, and statistical evaluations of probabilistic systems. What we need is someone who can design and operate robust evaluation frameworks, partner with scientists and engineers, and ensure we can confidently answer questions like: "Did this change improve or degrade quality, safety, or user outcomes?"

What You'll Build

Evaluation & Quality Pipelines

  • Build data evaluation pipelines that collect production conversations and agent interactions
  • Reconstruct full sessions from traces, logs, recordings, and transcripts
  • Apply labeling and scoring using human feedback signals (surveys, sentiment, outcomes) and automated evaluators (e.g., LLMasjudge)

Continuous Quality & Safety Benchmarking

  • Own weekly and ondemand automated evaluation runs against staging and production
  • Define benchmarks that track accuracy, reliability, and safetyrelated signals
  • Produce trend dashboards that clearly answer: "Did this deploy change quality or risk?"

Unified Evaluation Framework

  • Design and extend a standardized evaluation framework that supports multiple agent types and workflows
  • Translate highlevel product expectations into concrete success criteria and metrics
  • Ensure new agents and features can be evaluated consistently with minimal friction

Self Service Evaluation Tooling

  • Build APIs and internal tools so data scientists and engineers can go from "interesting scenario" to "included in the eval suite" quickly
  • Enable scenario curation, dataset management, and eval execution without deep infrastructure knowledge

Experiment Tracking & Visibility

  • Provide shared visibility into prompt, model, and agent experiments
  • Enable reproducibility and comparison across runs so teams can build on each other's work instead of operating in silos

Position Responsibilities:

Data Engineering

  • Build and maintain ETL pipelines for heterogeneous data sources (traces, logs, transcripts, user feedback)
  • Implement complex data stitching and session reconstruction logic
  • Manage dataset versioning, provenance, and lifecycle

Platform & Observability

  • Develop dashboards and monitoring tools for AI quality metrics
  • Integrate evaluations into CI/CD pipelines for scheduled and gated runs
  • Implement alerting on quality and safety signals, not just infrastructure health

AI / ML Evaluation Tooling

  • Apply and extend LLMasjudge evaluation patterns
  • Design metrics and scoring approaches suitable for stochastic, nondeterministic systems
  • Use tools like LangSmith to track runs, traces, experiments, and evaluation results

Collaboration

  • Partner closely with data science, engineering, and product teams
  • Translate between research goals, product intent, and engineering constraints
  • Help define what "good" looks like for AI behavior in production
  • Advocate for strong developer experience and usability in the tools you build

Required Qualifications

  • 4+ years of experience in data engineering, ML engineering, or software engineering
  • Bachelor's or Master's degree in Computer Science, Machine Learning, or a related quantitative field
    * Strong proficiency in Python
    * Experience building and maintaining production data pipelines
    * Strong SQL skills
    * Experience working with at least one cloud platform (AWS preferred)

NicetoHaves

  • Prior work on LLM or agent evaluation infrastructure
  • Familiarity with designing metrics for safety, reliability, or quality in AI systems
  • Experience with voice or callcenter data (audio, transcripts, sentiment)
  • Experience with browser automation tools (e.g., Playwright) for endtoend evals
  • Deep SQL expertise
New York, NY Salary Range
$161,600—$200,000 USD
Denver, CO Salary Range
$148,400—$185,000 USD
Charlotte, NC Salary Range
$134,800—$168,500 USD

All employees are responsible for adherence to the Judi Health Code of Conduct including the reporting of non-compliance. This position description is designed to be flexible, allowing management the opportunity to assign or reassign duties and responsibilities as needed to best meet organizational goals.

We provide equal employment opportunities to all employees and applicants for employment and prohibit discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, medical condition, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

By submitting an application, you agree to the retention of your personal data for consideration for a future position at Judi Health. More details about Judi Health's privacy practices can be found athttps://www.judi.health/legal/privacy-policy.