1

Coding Evaluator Jobs (NOW HIRING)

Evaluator **Company:** Adult & Child Health **Location:** *Indianapolis/Greenwood* **Employment ... Adhere to operational guidelines, organizational policies, and professional codes of ethics.* Meet ...

Evaluator Company: Adult & Child Health Location: Indianapolis/Greenwood Employment Type ... Adhere to operational guidelines, organizational policies, and professional codes of ethics. * Meet ...

Evaluators may also provide brief therapeutic interventions to support clients during the ... Adhere to operational guidelines, organizational policies, and professional codes of ethics.Meet ...

Evaluator **Company:** Adult & Child Health **Location:** *Indianapolis/Greenwood* **Employment ... Adhere to operational guidelines, organizational policies, and professional codes of ethics.* Meet ...

Evaluator Company: Adult & Child Health Location: Indianapolis/Greenwood Employment Type ... Adhere to operational guidelines, organizational policies, and professional codes of ethics. * Meet ...

Quality Evaluator

Novi, MI · On-site

$56K - $84K/yr

The Quality Evaluator is responsible for designing, implementing, and overseeing evaluation ... code of ethics, applicable federal, state, and local laws and regulations, HIPAA standards, and ...

next page

Showing results 1-20

Coding Evaluator information

See salary details

$29.5K

$65.5K

$106.5K

How much do coding evaluator jobs pay per year?

As of Sep 13, 2026, the average yearly pay for coding evaluator in the United States is $65,471.00, according to ZipRecruiter salary data. Most workers in this role earn between $44,500.00 and $79,500.00 per year, depending on experience, location, and employer.

What are popular job titles related to Coding Evaluator jobs?

For Coding Evaluator jobs, the most frequently searched job titles are:

Infographic showing various Coding Evaluator job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 83% Full Time, 10% Part Time, and 6% Contract. Highlights an 78% Physical, 4% Hybrid, and 18% Remote job distribution, with an average salary of $65,471 per year, or $31.5 per hour.

Principal Coding Annotator / LLM Evaluation Engineer

Alcolu, SC • On-site, Remote

$75 - $90/hr

Full-time

Posted 10 days ago


Job description

Company
Braintrust is a global talent network that connects top independent professionals with leading companies for high-quality, flexible work. We help organizations hire skilled talent faster while giving professionals access to vetted opportunities with innovative teams. Job description

This is a contracting engagement - initially 6 months - with potential for long term engagement.

Location: Paris or London-based preferred; alternatively Europe remote for strong candidates


We are building and evaluating state-of-the-art large language models (LLMs) and are looking for experienced software engineers to join our evaluation and annotation team. This role sits at the intersection of real-world software engineering, model evaluation, and applied AI, and is critical to improving model reliability, reasoning, and code quality.

You will design challenging coding tasks, evaluate model outputs against rigorous benchmarks, identify failure modes, and contribute to reinforcement learning and model improvement workflows.

This is not a junior annotation role. We are looking for practitioners with deep hands-on coding experience who can think like both an engineer and an evaluator.

What You’ll Do
  • Evaluate coding tasks involving software vulnerabilities, exploit verification, and security patches.
  • Create high-quality coding prompts and reference answers (benchmark-style, e.g. SWE-Bench-like problems).
  • Evaluate LLM outputs for code generation, refactoring, debugging, and implementation tasks.
  • Identify and document model failures, edge cases, and reasoning gaps.
  • Perform head-to-head evaluations between private LLMs (Mistral-based) and leading external models.
  • Build or configure coding environments to support evaluation and reinforcement learning (RL).
  • Follow detailed annotation and evaluation guidelines with high consistency.
What We’re Looking For
  • 5+ years of professional software development experience.
  • Strong Python skills (required).
  • Knowledge of at least one additional programming language (bonus).
  • Experience with professional code review, coding annotation, LLM/code evaluation, or benchmark design is a plus, but not required.
  • Hands-on experience with vulnerability research, exploit reproduction or verification, or implementing, backporting, or validating security patches.
  • Proven ability to apply structured evaluation criteria and write clear technical feedback.
  • Fluent in English (written and spoken).
  • Team lead or mentoring experience is a strong plus.
Why This Role
  • Work hands-on with cutting-edge LLMs.
  • Apply real-world engineering judgment to model evaluation and improvement.
  • High-impact, technical work with a focused, senior team.