... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
... language model post-training and evaluation, to produce the expert reasoning and evaluation data ... annotator would apply them the same way Why join Cobalt AI: * Advance frontier AI where it counts.
Instrument model quality monitoring in production -- detecting degradation across language pairs ... Feed evaluation insights back into data acquisition and model training priorities -- identifying ...
Instrument model quality monitoring in production -- detecting degradation across language pairs ... Feed evaluation insights back into data acquisition and model training priorities -- identifying ...
Instrument model quality monitoring in production -- detecting degradation across language pairs ... Feed evaluation insights back into data acquisition and model training priorities -- identifying ...
Instrument model quality monitoring in production -- detecting degradation across language pairs ... Feed evaluation insights back into data acquisition and model training priorities -- identifying ...
Phonetician
Menlo Park, CA ยท On-site +1
$40 - $45/hr
Responsibilities Perform narrow and broad phonetic transcriptions of speech data using the ... annotator reliability exercises Required Qualifications Bachelor's degree (or higher) in ...
Phonetician
Menlo Park, CA ยท On-site +1
$40 - $45/hr
Responsibilities Perform narrow and broad phonetic transcriptions of speech data using the ... annotator reliability exercises Required Qualifications Bachelor's degree (or higher) in ...
NLP Scientist
South San Francisco, CA ยท On-site
$75.23 - $83.23/hr
Strong Python production engineering with modern natural language processing frameworks ... annotator agreement. * Reports error rates by direction, not only aggregate accuracy, false ...
Quick apply
NLP Scientist
South San Francisco, CA ยท On-site
$75.23 - $83.23/hr
Strong Python production engineering with modern natural language processing frameworks ... annotator agreement. * Reports error rates by direction, not only aggregate accuracy, false ...
NLP Scientist
South San Francisco, CA ยท On-site
$75.23 - $83.23/hr
Strong Python production engineering with modern natural language processing frameworks ... annotator agreement. * Reports error rates by direction, not only aggregate accuracy, false ...
NLP Scientist
South San Francisco, CA ยท On-site
$75.23 - $83.23/hr
Strong Python production engineering with modern natural language processing frameworks ... annotator agreement. * Reports error rates by direction, not only aggregate accuracy, false ...
Producing concise summaries of longer pieces of text or data. * Converting spoken language or audio content into written text. * Converting text or spoken language from one language to another.
Producing concise summaries of longer pieces of text or data. * Converting spoken language or audio content into written text. * Converting text or spoken language from one language to another.
Producing concise summaries of longer pieces of text or data. * Converting spoken language or audio content into written text. * Converting text or spoken language from one language to another.
Producing concise summaries of longer pieces of text or data. * Converting spoken language or audio content into written text. * Converting text or spoken language from one language to another.
Hands-on experience fine-tuning large language models (open-weight models such as Llama, Mistral ... human data pipelines - annotation workflows, quality assurance methodology, inter-annotator ...
Hands-on experience fine-tuning large language models (open-weight models such as Llama, Mistral ... human data pipelines - annotation workflows, quality assurance methodology, inter-annotator ...
Resource Employee - Video Annotator
Los Angeles, CA ยท On-site
$50 - $100/hr
The data gathered by both research participants and resource employees will be used to train AI ... language is not English, and especially people who are formerly incarcerated and people who have ...
Resource Employee - Video Annotator
Los Angeles, CA ยท On-site
$50 - $100/hr
The data gathered by both research participants and resource employees will be used to train AI ... language is not English, and especially people who are formerly incarcerated and people who have ...
Resource Employee - Video Annotator
Los Angeles, CA ยท On-site
$30/hr
The data gathered by both research participants and resource employees will be used to train AI ... language is not English, and especially people who are formerly incarcerated and people who have ...
Resource Employee - Video Annotator
Los Angeles, CA ยท On-site
$30/hr
The data gathered by both research participants and resource employees will be used to train AI ... language is not English, and especially people who are formerly incarcerated and people who have ...
The data gathered by both research participants and resource employees will be used to train AI ... language is not English, and especially people who are formerly incarcerated and people who have ...
The data gathered by both research participants and resource employees will be used to train AI ... language is not English, and especially people who are formerly incarcerated and people who have ...
Language Data Annotator information
What is the difference between Language Data Annotator vs Transcription Specialist?
| Aspect | Language Data Annotator | Transcription Specialist |
|---|---|---|
| Credentials | Basic computer skills, attention to detail | Typing proficiency, language skills |
| Work Environment | Data labeling platforms, remote or office | Audio/video transcription, remote or office |
| Industry Usage | AI training, machine learning datasets | Media, legal, medical transcription |
| Search Intent | Compare roles in data annotation | Compare roles in transcription |
The main difference is that Language Data Annotators focus on labeling and categorizing language data for AI models, while Transcription Specialists convert audio or video content into written text. Both roles require language skills and attention to detail but serve different purposes within the language processing industry.
What cities are hiring for Language Data Annotator jobs?
Cities with the most Language Data Annotator job openings:
What states have the most Language Data Annotator jobs?
States with the most job openings for Language Data Annotator jobs include:
What are popular job titles related to Language Data Annotator jobs?
For Language Data Annotator jobs, the most frequently searched job titles are:

LLM Post-Training and Evaluation Researcher (Contract)
Santa Rosa, CA โข On-site
Other
This job post hasย expired today.ย Applications are no longer accepted.
Key responsibilities
Produce written reasoning traces and expert reference answers on technical problems.
Evaluate and rank model outputs on technical questions, articulating differences between responses.
Design rubrics, reward criteria, and partial-credit schemes for multistep tasks.
Job description
About the role:
Cobalt is seeking researchers and engineers with direct experience in language model post-training and evaluation, to produce the expert reasoning and evaluation data frontier labs use to improve model behavior.
This opportunity is suited to people who have worked on the parts of the stack closest to how a model actually behaves: supervised fine-tuning, preference optimization and RLHF, reward modeling, inference-time reasoning methods, and the design of evaluations that hold up. You may have done this in a lab, in industry, or in serious open-source work.
You do not need prior experience in data annotation. What matters is that you understand why models fail in the ways they do, and that you can write the kind of data and criteria that fix it.
What you'll do:
Depending on the project, you may:
- Produce written reasoning traces and expert reference answers on hard technical problems, at the standard of quality a post-training set requires rather than merely correct answers
- Author evaluation items and benchmark tasks with verifiable success criteria, including cases specifically designed to separate genuine capability from pattern matching
- Evaluate and rank model outputs on technical questions, articulating precisely what separates a strong response from one that is fluent but subtly wrong
- Design rubrics, reward criteria, and partial-credit schemes for multistep tasks, and identify where a criterion would be gameable or would reward the wrong behavior
- Classify observed failures into a consistent taxonomy, and assess whether a stated conclusion is supported by the underlying reasoning
Projects follow their own guidelines, formatting conventions, and quality standards, and you will work with feedback from reviewers and lab research teams.
Required qualifcations
- Direct experience with language model post-training or evaluation, such as supervised fine-tuning, preference optimization or RLHF, reward modeling, or benchmark and eval design, whether in research or industry
- A PhD in a quantitative discipline, or equivalent depth demonstrated through published work, open-source contributions, or production systems
- Strong coding ability in Python, and working command of at least one deep learning framework
- Understanding of common failure modes in current models, including reward hacking, sycophancy, and answers that are right for the wrong reasons
- Ability to explain each step of your reasoning clearly in writing, and to specify criteria precisely enough that another annotator would apply them the same way
Why join Cobalt AI:
- Advance frontier AI where it counts. Apply your expertise to data that frontier labs cannot obtain any other way, where your reasoning directly shapes how the next generation of models works through technical problems.
- Grow professionally. Expand your influence through evaluation projects, advisory roles, and research collaborations, while developing a working understanding of how frontier models are trained and assessed.
- Work with a top-tier network. Collaborate with researchers and engineers from leading institutions and labs on high-impact, flexible work.
- Set your own schedule. Flexible 10 to 40 hour weeks that fit around your existing work and your life.
- Competitive pay. Rates vary by project and are determined by a number of factors, including scope, skillset, and experience.