LLM-as-a-judge) to support the humans-in-the-loop (HITL) work of the broader Editorial team. They will lead and drive the requirements behind data annotation tasks and tooling, writing intuitive ...
LLM-as-a-judge) to support the humans-in-the-loop (HITL) work of the broader Editorial team. They will lead and drive the requirements behind data annotation tasks and tooling, writing intuitive ...
... data annotation and rubric-based scoring • Prior work in trust and safety, content moderation, QA, or security research • Subject matter expertise in any high-risk domain (cybersecurity ...
... data annotation and rubric-based scoring • Prior work in trust and safety, content moderation, QA, or security research • Subject matter expertise in any high-risk domain (cybersecurity ...
LLM-as-a-judge) to support the humans-in-the-loop (HITL) work of the broader Editorial team. They will lead and drive the requirements behind data annotation tasks and tooling, writing intuitive ...
LLM-as-a-judge) to support the humans-in-the-loop (HITL) work of the broader Editorial team. They will lead and drive the requirements behind data annotation tasks and tooling, writing intuitive ...
LLM-as-a-judge) to support the humans-in-the-loop (HITL) work of the broader Editorial team. They will lead and drive the requirements behind data annotation tasks and tooling, writing intuitive ...
LLM-as-a-judge) to support the humans-in-the-loop (HITL) work of the broader Editorial team. They will lead and drive the requirements behind data annotation tasks and tooling, writing intuitive ...
... annotation modalities, and cold start conditions Build and maintain calibration frameworks that keep LLM evaluators anchored to human judgment over time Develop anomaly detection systems that surface ...
... annotation modalities, and cold start conditions Build and maintain calibration frameworks that keep LLM evaluators anchored to human judgment over time Develop anomaly detection systems that surface ...
Linguistic Managed Services Support (Tamil, Marathi, Egyptian Arabic and Spanish) !! Seatle, WA (O
Seattle, WA · On-site
... LLM language programs. Requirements: * Native/Near-native proficiency in Tamil, Marathi, or ... Experience with language annotation, data labeling, or linguistic QA is a plus.
Linguistic Managed Services Support (Tamil, Marathi, Egyptian Arabic and Spanish) !! Seatle, WA (O
Seattle, WA · On-site
... LLM language programs. Requirements: * Native/Near-native proficiency in Tamil, Marathi, or ... Experience with language annotation, data labeling, or linguistic QA is a plus.
AI Red Teamer, LLM Generalist
Seattle, WA · On-site
$32 - $95/hr
Experience working with LLM APIs or evaluation tooling ... Comfort with structured data annotation and rubric-based scoring * Prior work in trust and safety ...
Quick apply
AI Red Teamer, LLM Generalist
Seattle, WA · On-site
$32 - $95/hr
Experience working with LLM APIs or evaluation tooling ... Comfort with structured data annotation and rubric-based scoring * Prior work in trust and safety ...
Red Teaming Expert
Seattle, WA · On-site
$30 - $40/hr
LLM red teaming or AI safety evaluation * Trust & safety, content moderation, or policy enforcement * AI/ML evaluation, annotation, or QA workflows * Conversational analysis or behavioral risk ...
Red Teaming Expert
Seattle, WA · On-site
$30 - $40/hr
LLM red teaming or AI safety evaluation * Trust & safety, content moderation, or policy enforcement * AI/ML evaluation, annotation, or QA workflows * Conversational analysis or behavioral risk ...
Familiarity with prompt engineering, tool-use patterns, and LLM model behavior. Experience ... dataset generation, annotation workflows, data validation, or high-throughput processing)
Familiarity with prompt engineering, tool-use patterns, and LLM model behavior. Experience ... dataset generation, annotation workflows, data validation, or high-throughput processing)
Familiarity with prompt engineering, tool-use patterns, and LLM model behavior. Experience ... dataset generation, annotation workflows, data validation, or high-throughput processing)
Familiarity with prompt engineering, tool-use patterns, and LLM model behavior. Experience ... dataset generation, annotation workflows, data validation, or high-throughput processing)
Experience with AI training, data annotation, large language models, prompt/response evaluation, or rubric-based LLM QA is a strong plus. Key responsibilities * Quality monitoring: Spot-check Kannada ...
Experience with AI training, data annotation, large language models, prompt/response evaluation, or rubric-based LLM QA is a strong plus. Key responsibilities * Quality monitoring: Spot-check Kannada ...
Experience with AI training, data annotation, large language models, prompt/response evaluation, or rubric-based LLM QA is a strong plus. Key responsibilities * Quality monitoring: Spot-check ...
Experience with AI training, data annotation, large language models, prompt/response evaluation, or rubric-based LLM QA is a strong plus. Key responsibilities * Quality monitoring: Spot-check ...
Experience with AI training, data annotation, large language models, prompt/response evaluation, or rubric-based LLM QA is a strong plus. Key responsibilities * Quality monitoring: Spot-check Telugu ...
Experience with AI training, data annotation, large language models, prompt/response evaluation, or rubric-based LLM QA is a strong plus. Key responsibilities * Quality monitoring: Spot-check Telugu ...
Senior Manager, Machine Learning Engineering
$200K - $250K/yr
Build scalable data engineering pipelines and automated annotation workflows (LLM-in-the-loop) to reduce reliance on manual labeling and accelerate model iteration * Own the MLOps lifecycle ...
Quick apply
Senior Manager, Machine Learning Engineering
$200K - $250K/yr
Build scalable data engineering pipelines and automated annotation workflows (LLM-in-the-loop) to reduce reliance on manual labeling and accelerate model iteration * Own the MLOps lifecycle ...
Senior Manager, Machine Learning Engineering
$200K - $250K/yr
Build scalable data engineering pipelines and automated annotation workflows (LLM-in-the-loop) to reduce reliance on manual labeling and accelerate model iteration * Own the MLOps lifecycle ...
Senior Manager, Machine Learning Engineering
$200K - $250K/yr
Build scalable data engineering pipelines and automated annotation workflows (LLM-in-the-loop) to reduce reliance on manual labeling and accelerate model iteration * Own the MLOps lifecycle ...
Senior Manager, Machine Learning Engineering
Seattle, WA · On-site
$200K - $250K/yr
Build scalable data engineering pipelines and automated annotation workflows (LLM-in-the-loop) to reduce reliance on manual labeling and accelerate model iteration * Own the MLOps lifecycle ...
Senior Manager, Machine Learning Engineering
Seattle, WA · On-site
$200K - $250K/yr
Build scalable data engineering pipelines and automated annotation workflows (LLM-in-the-loop) to reduce reliance on manual labeling and accelerate model iteration * Own the MLOps lifecycle ...
Wellness & Nutrition Content Expert (Contract)
Seattle, WA · On-site
$30 - $40/hr
... annotation for LLM builders. Our reviewers have specialization in behavioral analysis, conversational design, mental health, psychiatry, social services and clinical trial settings. About the Role ...
Wellness & Nutrition Content Expert (Contract)
Seattle, WA · On-site
$30 - $40/hr
... annotation for LLM builders. Our reviewers have specialization in behavioral analysis, conversational design, mental health, psychiatry, social services and clinical trial settings. About the Role ...
... annotation for LLM builders. Our reviewers have specialization in behavioral analysis, conversational design, mental health, psychiatry, social services and clinical trial settings. About the Role ...
Quick apply
... annotation for LLM builders. Our reviewers have specialization in behavioral analysis, conversational design, mental health, psychiatry, social services and clinical trial settings. About the Role ...
Wellness & Nutrition Content Expert (Contract)
Seattle, WA · Remote
$30 - $40/hr
... annotation for LLM builders. Our reviewers have specialization in behavioral analysis, conversational design, mental health, psychiatry, social services and clinical trial settings. About the Role ...
New
Quick apply
Wellness & Nutrition Content Expert (Contract)
Seattle, WA · Remote
$30 - $40/hr
... annotation for LLM builders. Our reviewers have specialization in behavioral analysis, conversational design, mental health, psychiatry, social services and clinical trial settings. About the Role ...
New
Wellness & Nutrition Content Expert (Contract)
Seattle, WA · On-site +1
$30 - $40/hr
... annotation for LLM builders. Our reviewers have specialization in behavioral analysis, conversational design, mental health, psychiatry, social services and clinical trial settings. About the Role ...
Wellness & Nutrition Content Expert (Contract)
Seattle, WA · On-site +1
$30 - $40/hr
... annotation for LLM builders. Our reviewers have specialization in behavioral analysis, conversational design, mental health, psychiatry, social services and clinical trial settings. About the Role ...
Llm Annotation information
See Seattle, WA salary details
$12.5K - $15.7K
0% of jobs
$15.7K - $18.8K
0% of jobs
$18.8K - $22K
0% of jobs
$22K - $25.1K
0% of jobs
$25.1K - $28.3K
0% of jobs
$28.3K - $31.5K
0% of jobs
$31.5K - $34.6K
0% of jobs
$34.6K - $37.8K
0% of jobs
$37.8K - $40.9K
0% of jobs
$40.9K - $44.1K
0% of jobs
$44.9K is the 25th percentile. Wages below this are outliers.
$44.1K - $47.2K
100% of jobs
$12.5K
$47.2K
How much do llm annotation jobs pay per year?
Which 5 jobs will survive AI?
How much do AI annotators make?
Are data annotations still hiring?
What is an LLM annotator?
What is the difference between Llm Annotation vs Data Labeler?
| Aspect | Llm Annotation | Data Labeler |
|---|---|---|
| Required Credentials | Basic computer skills, sometimes familiarity with AI tools | Basic skills, often on-the-job training |
| Work Environment | Remote or office-based, tech-focused | Remote or on-site, varied industries |
| Industry Usage | AI, machine learning, NLP projects | Various industries including marketing, healthcare, and tech |
| Search & Comparison Intent | Understanding roles in AI data preparation | General data labeling tasks |
In summary, Llm Annotation involves specialized annotation for large language models, often requiring familiarity with AI tools, while Data Labeler is a broader role focused on labeling data across multiple industries with minimal technical requirements.
What is LLM annotation?
What are the key skills and qualifications needed to thrive as an LLM Annotation Specialist, and why are they important?
What are some common challenges faced by LLM Annotation specialists, and how can they be addressed?
Amazon rating
7.4
Based on 7,045 frontline employees who took The Breakroom Quiz
6th of 39 rated national retailers
Job description
The Conversational Shopping team is looking for a Language Engineer to drive efficiencies and innovation in its efforts to deliver a seamless, fluent, and engaging experience for AI-assisted shopping. This is an opportunity to join the high-performing team behind Amazon's Generative AI shopping initiatives such as Rufus AI, Amazon's Conversational Shopping assistant. Our objective is to make it easy for customers worldwide to find and discover the best products, meet their personalized needs with product research, providing comparisons and recommendations, answering specific product questions, and more.
This role is inherently high-visibility and highly cross-functional, requiring collaboration and influence across global product, design, science, and engineering teams.
We are looking for candidates who are passionate about the intersection of language and technology and who are keen to use their technical abilities to develop automated, scalable solutions to questions in the Large Language Model (LLM) space. Applying a combination of expertise in LLMs, coding and linguistics (i.e., semantics, syntax, pragmatics), they will overcome complex problems in natural language processing (NLP), language understanding and automated AI evaluations.
In this role within the Editorial team, they will act as one of the driving forces behind our evaluation-driven product development strategy. They will design processes to facilitate the production of high quality editorial data which will allow us to evaluate and improve the Shopping AI experience in different languages
To do so, they will be tasked with the creation and development of LLM-assisted editorial tools, automated verification scripts and automated annotations (e.g. LLM-as-a-judge) to support the humans-in-the-loop (HITL) work of the broader Editorial team. They will lead and drive the requirements behind data annotation tasks and tooling, writing intuitive annotation guidelines and guiding the creation of the tools adapted to these workflows
They will employ their data processing and analysis skills to track team productivity and measure output quality. They will work in close collaboration with other Language Engineers, AI Editors, Product Managers, Applied Scientists and Software Engineers on initiatives that drive editorial quality, speed and consistency. By creating and synthesizing quality metrics, they will also guide Conversational Shopping teams in delivering both internal stakeholder requirements and achieve the desired Amazon customer outcomes.
This role requires strong analytical and technical skills as well as experience in language technology to help us measure, analyze and solve complex problems
They should have experience in creating technical solutions for automating and processing data workflows at scale and have the ability to do so while upholding the highest linguistic quality standards. They should also have exceptional writing and communication skills with the ability to interface between both technical and non-technical teams.
Key job responsibilities
* Produce, process and manipulate different types of language data, analyze, and provide efficient solutions
* Automate operations and perform data analysis using coding/scripting language (e.g. Python)
* Develop LLM-assisted workflows and annotations solutions (e.g
LLM-as-a-judge) to support Human-in-the-loop evaluations
* Design and lead editorial data production/collection by defining scope with internal customer teams
* Define clear editorial workflows (SOPs) to meet or exceed the quality bar
* Adopt and design control mechanisms, metrics and methodologies for editorial and annotation quality
* Maximize productivity, process efficiency and quality through streamlined workflows, process standardization, documentation, audits and investigations on a periodic basis.
* Collaborate with editors, applied scientists, engineers, and product managers to deliver the optimal customer experience and define metrics, guidelines, and workflows to continue doing so
* Establish processes and mechanisms to onboard and train editors on an ongoing basis.
* Handle work prioritization and deliver based on business priorities.
* Be flexible in changes to conventions deployed in response to customers' requests and change workflows accordingly.
A day in the life
- Build an LLM-powered judge that automatically scores thousands of AI responses
- Write a clear evaluation guideline, then pair with a PM-T and Scientist to validate it captures what "good" actually looks like
- Debug a Python pipeline that processes multilingual annotation data, spot a pattern in the errors, and ship a fix
- Join a cross-functional sync to align on quality metrics for an upcoming feature launch
- Analyze evaluation results to surface insights that shape what the product team prioritizes next
You'll split your time between hands-on technical work (code, prompts, data) and collaborative problem-solving with editors, engineers, and PMs.
About the team
We're the language technology team within Amazon's conversational shopping organization. We build and operate LLM-as-a-Judge systems that automatically measure response quality across every customer experience, develop agentic evaluation architectures that resolve cases single-hop judges can't, and create the tooling and automation that let a small team evaluate millions of AI responses at scale. You'll work alongside language engineers, AI editors, data scientists, and product managers, collaborating cross-functionally with applied scientists and software engineers to ship judges, define quality standards, and turn evaluation data into product decisions for customers around the world.
About Amazon
Sourced by ZipRecruiter
Amazon.com, Inc., commonly known as Amazon, is an American multinational technology company. It was founded by Jeff Bezos in 1994 and initially started as an online marketplace for books. Since then, Amazon has expanded its operations and become one of the largest e-commerce companies in the world. Amazon's primary business is its online retail platform, where customers can purchase a vast array of products, including electronics, clothing, books, home goods, and much more. The company offers a convenient and user-friendly shopping experience, with features such as fast shipping, customer reviews, and personalized recommendations. In addition to its e-commerce platform, Amazon has diversified its business into various other areas. One of its notable ventures is Amazon Web Services (AWS), a comprehensive cloud computing platform that provides services such as storage, compute power, and database management to individuals and businesses. AWS has become a leader in the cloud computing industry, powering many websites and applications worldwide. Amazon has also developed its own consumer electronics, including the popular Amazon Kindle e-reader, Fire tablets, Fire TV streaming devices, and the Alexa-powered Echo smart speakers. The Alexa voice assistant, integrated into these devices, allows users to interact with their devices using voice commands, perform tasks, and access information. Furthermore, Amazon has expanded into media and entertainment. It operates Prime Video, a streaming service that offers a wide range of movies, TV shows, and original content. Amazon Music provides a platform for streaming and purchasing digital music, while Audible offers audiobooks and other audio content. The company's commitment to customer satisfaction and convenience is demonstrated by its membership program, Amazon Prime. Prime members receive various benefits, including free two-day shipping, access to streaming services, exclusive deals, and more.
Industry
It services, book publishers, retail, real estate and computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Seattle, WA, US