1

Language Data Annotator Jobs (NOW HIRING)

Instrument model quality monitoring in production -- detecting degradation across language pairs ... Feed evaluation insights back into data acquisition and model training priorities -- identifying ...

Phonetician

Menlo Park, CA ยท On-site +1

$40 - $45/hr

Responsibilities Perform narrow and broad phonetic transcriptions of speech data using the ... annotator reliability exercises Required Qualifications Bachelor's degree (or higher) in ...

NLP Scientist

South San Francisco, CA ยท On-site

$75.23 - $83.23/hr

Strong Python production engineering with modern natural language processing frameworks ... annotator agreement. * Reports error rates by direction, not only aggregate accuracy, false ...

Showing results 41-59

Language Data Annotator information

What is the difference between Language Data Annotator vs Transcription Specialist?

AspectLanguage Data AnnotatorTranscription Specialist
CredentialsBasic computer skills, attention to detailTyping proficiency, language skills
Work EnvironmentData labeling platforms, remote or officeAudio/video transcription, remote or office
Industry UsageAI training, machine learning datasetsMedia, legal, medical transcription
Search IntentCompare roles in data annotationCompare roles in transcription

The main difference is that Language Data Annotators focus on labeling and categorizing language data for AI models, while Transcription Specialists convert audio or video content into written text. Both roles require language skills and attention to detail but serve different purposes within the language processing industry.

What cities are hiring for Language Data Annotator jobs?

Cities with the most Language Data Annotator job openings:

What states have the most Language Data Annotator jobs?

States with the most job openings for Language Data Annotator jobs include:

What are popular job titles related to Language Data Annotator jobs?

For Language Data Annotator jobs, the most frequently searched job titles are:

Infographic showing various Language Data Annotator job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 1% As Needed, 84% Full Time, 11% Part Time, and 3% Contract. Highlights an 85% Physical, 3% Hybrid, and 12% Remote job distribution.

LLM Post-Training and Evaluation Researcher (Contract)

Santa Rosa, CA โ€ข On-site

Other

This job post hasย expired today.ย Applications are no longer accepted.


Key responsibilities

  • Produce written reasoning traces and expert reference answers on technical problems.

  • Evaluate and rank model outputs on technical questions, articulating differences between responses.

  • Design rubrics, reward criteria, and partial-credit schemes for multistep tasks.


Job description

About the role:

Cobalt is seeking researchers and engineers with direct experience in language model post-training and evaluation, to produce the expert reasoning and evaluation data frontier labs use to improve model behavior.

This opportunity is suited to people who have worked on the parts of the stack closest to how a model actually behaves: supervised fine-tuning, preference optimization and RLHF, reward modeling, inference-time reasoning methods, and the design of evaluations that hold up. You may have done this in a lab, in industry, or in serious open-source work.

You do not need prior experience in data annotation. What matters is that you understand why models fail in the ways they do, and that you can write the kind of data and criteria that fix it.


What you'll do:

Depending on the project, you may:

  • Produce written reasoning traces and expert reference answers on hard technical problems, at the standard of quality a post-training set requires rather than merely correct answers
  • Author evaluation items and benchmark tasks with verifiable success criteria, including cases specifically designed to separate genuine capability from pattern matching
  • Evaluate and rank model outputs on technical questions, articulating precisely what separates a strong response from one that is fluent but subtly wrong
  • Design rubrics, reward criteria, and partial-credit schemes for multistep tasks, and identify where a criterion would be gameable or would reward the wrong behavior
  • Classify observed failures into a consistent taxonomy, and assess whether a stated conclusion is supported by the underlying reasoning

Projects follow their own guidelines, formatting conventions, and quality standards, and you will work with feedback from reviewers and lab research teams.


Required qualifcations

  • Direct experience with language model post-training or evaluation, such as supervised fine-tuning, preference optimization or RLHF, reward modeling, or benchmark and eval design, whether in research or industry
  • A PhD in a quantitative discipline, or equivalent depth demonstrated through published work, open-source contributions, or production systems
  • Strong coding ability in Python, and working command of at least one deep learning framework
  • Understanding of common failure modes in current models, including reward hacking, sycophancy, and answers that are right for the wrong reasons
  • Ability to explain each step of your reasoning clearly in writing, and to specify criteria precisely enough that another annotator would apply them the same way


Why join Cobalt AI:

  • Advance frontier AI where it counts. Apply your expertise to data that frontier labs cannot obtain any other way, where your reasoning directly shapes how the next generation of models works through technical problems.
  • Grow professionally. Expand your influence through evaluation projects, advisory roles, and research collaborations, while developing a working understanding of how frontier models are trained and assessed.
  • Work with a top-tier network. Collaborate with researchers and engineers from leading institutions and labs on high-impact, flexible work.
  • Set your own schedule. Flexible 10 to 40 hour weeks that fit around your existing work and your life.
  • Competitive pay. Rates vary by project and are determined by a number of factors, including scope, skillset, and experience.