1

Llm Annotation Jobs in Texas (NOW HIRING)

Responsibilities : โ€ข Develop, evaluate, and iterate on NLP and LLM-based systems, including text ... annotation guidelines and ensuring label quality. โ€ข Evaluate and apply the appropriate approach ...

Review existing models, datasets, annotation processes, and production use cases * Identify the highest-impact opportunities for NLP and LLM improvements * Define practical evaluation metrics for ...

Develop, evaluate, and iterate on NLP and LLM-based systems, including text classification ... Experience designing data annotation workflows, labeling guidelines, or label quality processes is ...

Data Scientist II

Austin, TX ยท On-site +1

Develop, evaluate, and iterate on NLP and LLM-based systems, including text classification ... Experience designing data annotation workflows, labeling guidelines, or label quality processes is ...

Delivery Lead

Austin, TX ยท Remote

$110K - $140K/yr

... LLM assistance Preferred Qualifications * 1+ year in AI data operations (RLHF, annotation, model evaluation) * STEM background or strong technical fluency * Python & REACT working knowledge

Delivery Lead

Dallas, TX ยท Remote

$110K - $140K/yr

... LLM assistance Preferred Qualifications * 1+ year in AI data operations (RLHF, annotation, model evaluation) * STEM background or strong technical fluency * Python & REACT working knowledge

Familiarity with annotation workflows and data quality frameworks * AI/LLM Evaluation: * Hands-on experience evaluating Large Language Models (LLMs) * Familiarity with evaluation frameworks such as ...

Familiarity with annotation workflows and data quality frameworks * AI/LLM Evaluation: * Hands-on experience evaluating Large Language Models (LLMs) * Familiarity with evaluation frameworks such as ...

Partner with data science/ML teams to frame evaluation metrics for LLM-powered features and ground truth/annotation workflows. * Validate data quality, lineage, and mappings across EHR, claims, and ...

next page

Showing results 1-20

Llm Annotation information

Which 5 jobs will survive AI?

Jobs involving LLM annotation, such as data annotators and labelers, are likely to persist as they require human judgment for complex or nuanced tasks. Roles that involve creative thinking, emotional intelligence, and strategic decision-making, like psychologists, teachers, healthcare professionals, and managers, are also expected to remain in demand despite AI advancements. These jobs often require skills that are difficult for AI to replicate fully.

How much do AI annotators make?

AI annotators, including those working as language model annotation specialists, typically earn between $12 and $20 per hour, depending on experience, location, and the complexity of the tasks. Some positions may offer hourly wages or project-based pay, with higher rates for specialized skills or advanced tools proficiency.

Are data annotations still hiring?

Data annotation roles, including those for large language models (LLMs), are currently in demand as companies continue to develop AI and machine learning systems. These jobs often require attention to detail and familiarity with annotation tools, and opportunities are available through various online platforms and companies expanding their AI teams.

What is an LLM annotator?

An LLM annotator is a person who labels and tags data to train large language models (LLMs). They review and annotate text data to improve model accuracy, often using specialized tools and following specific guidelines. This role requires attention to detail and understanding of language patterns.

What is the difference between Llm Annotation vs Data Labeler?

AspectLlm AnnotationData Labeler
Required CredentialsBasic computer skills, sometimes familiarity with AI toolsBasic skills, often on-the-job training
Work EnvironmentRemote or office-based, tech-focusedRemote or on-site, varied industries
Industry UsageAI, machine learning, NLP projectsVarious industries including marketing, healthcare, and tech
Search & Comparison IntentUnderstanding roles in AI data preparationGeneral data labeling tasks

In summary, Llm Annotation involves specialized annotation for large language models, often requiring familiarity with AI tools, while Data Labeler is a broader role focused on labeling data across multiple industries with minimal technical requirements.

What is LLM annotation?

LLM annotation refers to the process of labeling or tagging data specifically for training and evaluating large language models (LLMs) like GPT or BERT. Annotators read text and apply labels, correct errors, or provide feedback to help improve the model's understanding and performance. This work is crucial for supervised learning, as well-annotated datasets help LLMs better recognize patterns, context, and meaning in human language. LLM annotation can involve tasks such as sentiment analysis, named entity recognition, or instruction following. Annotators often use specialized platforms or tools to complete their tasks efficiently and accurately.

What are the key skills and qualifications needed to thrive as an LLM Annotation Specialist, and why are they important?

To thrive as an LLM Annotation Specialist, you need strong analytical skills, attention to detail, and a background in linguistics, computer science, or a related field. Familiarity with annotation platforms, natural language processing (NLP) tools, and data labeling systems is typically required. Excellent communication, critical thinking, and the ability to follow guidelines precisely are valuable soft skills for this role. These skills ensure high-quality, accurate data annotation, which directly impacts the performance and reliability of large language models.

What are some common challenges faced by LLM Annotation specialists, and how can they be addressed?

LLM Annotation specialists often encounter challenges such as interpreting ambiguous language data, maintaining annotation consistency across complex datasets, and keeping up with evolving guidelines. These can be addressed by participating in regular team syncs to clarify guidelines, using annotation tools with built-in quality checks, and collaborating closely with project leads and fellow annotators. Continuous learning and open communication help ensure high-quality, reliable data annotation and support professional growth within the AI and NLP fields.
What cities in Texas are hiring for Llm Annotation jobs? Cities in Texas with the most Llm Annotation job openings:
Infographic showing various Llm Annotation job openings in Texas as of July 2026, with employment types broken down into 1% As Needed, 52% Full Time, 45% Part Time, and 2% Contract. Highlights an 46% Physical, 1% Hybrid, and 53% Remote job distribution.

Applied Data Scientist, LLM Evaluation

Driver AI Inc.

Austin, TX โ€ข Remote

Other

Medical, Dental, Vision, Life, Retirement

Posted 4 days ago


Job description

Applied Data Scientist, LLM Evaluation Introduction

At Driver, we're building systems that turn source code into human language. The tech stack includes a core compiler-like engine, a heavily asynchronous/distributed backend server, and a frontend web application that provides a rich user experience.

About Driver

We're an early-stage startup backed by Y Combinator and Google Ventures that combines first principles technical approaches and applied LLM expertise to tackle context engineering at scale. Driver builds the context layer for employees and AI agents alike to use in developing software.

Working at Driver

Driver is an early-stage but fast-growing startup. As such, we take advantage of that which startups can excel: delivery speed, flexibility, and enjoying working with a small close-knit team.

Organizational and engineering values at Driver include first-principles thinking, correct by construction, writing things down, experimentation and iteration, pragmatism, commitment to effective communication and transparency, autonomy, and ambition.

Job Overview

Title: Applied Data Scientist, LLM Evaluation

Location: Remote or Austin, Tx

Our value is directly tied to the quality of our content at scale. The platform generates technical documentation across a complex, multi-stage pipeline - producing multiple content types at different levels of abstraction, from individual code elements up to high-level summaries. Today, changes to models, context strategies, or pipeline architecture are evaluated largely through manual review and intuition. There is no systematic way to answer: "Did this change make our output better, worse, or the same - and for which languages, repo sizes, and content types?"

This is a hard problem. LLM outputs are non-deterministic - identical inputs produce different outputs across runs, and small variations at early pipeline stages compound into meaningfully different end-user content downstream. Evaluating quality requires methodology that accounts for this: statistical reasoning over multiple runs, understanding of cascade effects through the pipeline, and rubrics that balance human judgment with automated signals.

This role builds the evaluation function from scratch. You'll define what "good" means for our generated content, build the infrastructure to measure it, and create the experimental framework that lets the team ship changes with confidence.

What You'll Do

You'll own the LLM evaluation strategy at Driver - from first principles to production infrastructure. This is a foundational role: you're not joining an existing eval team, you're building it. As the function matures, you'll seed and grow a team around it.

Define quality metrics and build evaluation datasets. Establish what "good" looks like for each content type across the pipeline. Build and curate gold-standard evaluation datasets across languages and repo archetypes (monorepos, microservices, libraries, applications). Design rubrics that capture accuracy, completeness, usefulness, and readability.

Build benchmarking and experimentation infrastructure. Create automated evaluation pipelines that score output against reference datasets. Instrument the content generation pipeline to support A/B comparisons - run the same codebase through two strategies and compare results. Build tooling for LLM-as-judge evaluation and regression detection. Integrate evaluation into CI so pipeline changes come with quality evidence.

Develop automated quality signals at scale. Build quality checks that flag degraded output without requiring human review of every document. Monitor content quality trends over time. Design sampling strategies for human review that maximize signal with minimal annotation effort.

Quantify tradeoffs and inform decisions. Run experiments on model selection, context strategies, and pipeline architecture changes. Quantify cost/quality/latency tradeoffs. Partner with the engineering team to turn evaluation insights into shipped improvements.

Qualifications

Education: Bachelor's, Master's, or PhD in Statistics, Machine Learning, Data Science, Computational Linguistics, or a related quantitative field.

Experience: Minimum 3 - 5 years in applied science, ML engineering, or data science roles with a focus on evaluation, NLP, or generative AI. 7+ years experience preferred.

Required Technical Skills

  • Strong statistical foundations: experimental design, hypothesis testing, confidence intervals, effect sizes, power analysis.
  • Experience designing and running evaluations for LLM or NLP systems - you've thought carefully about what "better" means when outputs are open-ended text.
  • Proficient in Python and the scientific/data stack (pandas, NumPy, scipy, sklearn).
  • Comfortable working in Jupyter notebooks for exploration and prototyping, and turning that work into automated pipelines.
  • Experience with LLM-as-judge approaches, inter-annotator agreement, and rubric design for subjective quality assessment.
  • Familiarity with the practical challenges of non-deterministic systems: variance decomposition, multi-run methodology, distinguishing signal from noise at scale.
  • Strong data storytelling - you can turn experiment results into clear recommendations that drive engineering and product decisions.

Preferred and Nice-to-Have Technical Skills

  • Experience with LLM APIs and prompt engineering across multiple providers.
  • Familiarity with evaluation frameworks (e.g., RAGAS, DeepEval, custom harnesses).
  • Experience building data pipelines or ETL workflows (Airflow, Dagster, or similar).
  • Comfort with SQL and working directly against production data stores.
  • Experience with visualization tools (Matplotlib, Plotly, Streamlit) for building internal dashboards and reports.
  • Background in code understanding, developer tools, or technical documentation.
  • Experience building or managing annotation pipelines and human evaluation workflows.
Benefits
  • Competitive Compensation Packages - Cash & Equity
  • Flexible Work Culture
  • Unlimited Time Off + 12 Paid Company Holidays
  • Insurance - Health, Dental, & Vision
  • Life Insurance & FSA Accounts
  • 401(k) Retirement Accounts - Traditional, Roth, or Both
  • Quarterly Team Offsites

Driver is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.