1

Ai Llm Data Labeling Jobs (NOW HIRING)

Seeking entry level data specialists/IT for a robotics/AI company. Fresh graduates are encouraged to apply. Responsibilities * Use proprietary annotation tools to label objects, poses, and ...

Senior AI/LLM Engineer

Irving, TX · On-site

$100K - $137K/yr

Title: Sr AI/LLM Engineer Location: Irvine, CA (Onsite) Duration: 6 months (possibility of an ... Faithfulness Relevance NDCG MRR Trace-level RAG evaluation (Langfuse) Data Engineering & ETL ...

We are seeking an AI/LLM Safety Engineer to join our AI team and take ownership of how safely our ... Lead structured red-teaming exercises covering jailbreaks, prompt injection, tool misuse, and data ...

Job Summary : Pyramid Consulting, Inc. is seeking talented AI/LLM Engineers for a long-term ... SQL and data platforms (Athena, Spark, or equivalent). • Strong software debugging skills ...

Remote AI/LLM Engineer

Odessa, FL · On-site

$90K - $122K/yr

Develop robust APIs and data models that enable seamless integration of AI services across multiple ... Strong understanding of LLM architecture, including context management, orchestration strategies ...

Data Labeling Specialist We are seeking a detail-oriented Data Labeling Specialist to join our ... Join a first-of-its-kind AI robotics company focused on bringing a general-purpose humanoid to life.

Seeking entry level data specialists/IT for a robotics/AI company. Fresh graduates are encouraged to apply. Responsibilities * Use proprietary annotation tools to label objects, poses, and ...

Data Labeling Specialist We are seeking a detail-oriented Data Labeling Specialist to join our ... Join a first-of-its-kind AI robotics company focused on bringing a general-purpose humanoid to life.

next page

Showing results 1-20

Ai Llm Data Labeling information

What is AI LLM data labeling?

AI LLM data labeling is the process of annotating or tagging data—such as text, images, or audio—to provide clear examples that help train large language models (LLMs) like GPT or BERT. This labeled data is essential for teaching models to understand context, intent, and meaning, which improves their performance on various tasks. Data labelers often follow specific guidelines to ensure consistency and accuracy, making their role critical in developing reliable AI systems.

What is the difference between Ai Llm Data Labeling vs Data Annotation Specialist?

AspectAi Llm Data LabelingData Annotation Specialist
CredentialsBasic technical skills, familiarity with labeling toolsSimilar technical skills, often with additional domain knowledge
Work EnvironmentData labeling platforms, remote or office settingsData annotation projects, remote or onsite
Industry UsageAI, machine learning, NLP projectsData preparation across various industries including AI

Ai Llm Data Labeling and Data Annotation Specialist roles both involve preparing data for machine learning models. However, Ai Llm Data Labeling typically focuses on labeling data specifically for large language models, requiring familiarity with NLP and AI tools. Data Annotation Specialists may work across broader data types and industries, with a focus on accurate data tagging. Both roles demand similar skills but differ in scope and application within AI projects.

What are the key skills and qualifications needed to thrive as an AI LLM Data Labeling Specialist, and why are they important?

To thrive as an AI LLM Data Labeling Specialist, you need keen attention to detail, strong analytical skills, and a foundational understanding of natural language processing concepts, often supported by familiarity with data annotation guidelines. Experience with labeling platforms (such as Labelbox or Prodigy), spreadsheet tools, and sometimes proficiency in scripting languages like Python is highly valued. Excellent communication, consistency, and critical thinking are crucial soft skills for interpreting ambiguous data and ensuring labeling accuracy. These skills and qualifications are vital for producing high-quality training data that directly impacts the performance and reliability of large language models.

What are some common challenges faced in AI LLM data labeling and how can they be managed?

One common challenge in AI LLM data labeling is ensuring consistency and accuracy when annotating large volumes of complex language data. Labelers often encounter ambiguous or context-dependent text, making it important to follow detailed guidelines and participate in regular calibration sessions with the team. Collaboration with data scientists and project managers is essential to clarify edge cases and refine labeling criteria. Proactively communicating questions and feedback helps maintain high-quality datasets, which are critical for training reliable language models.
More about Ai Llm Data Labeling jobs
What cities are hiring for Ai Llm Data Labeling jobs? Cities with the most Ai Llm Data Labeling job openings:
What states have the most Ai Llm Data Labeling jobs? States with the most job openings for Ai Llm Data Labeling jobs include:
What job categories do people searching Ai Llm Data Labeling jobs look for? The top searched job categories for Ai Llm Data Labeling jobs are:
Infographic showing various Ai Llm Data Labeling job openings in the United States as of July 2026, with employment types broken down into 75% Full Time, 22% Part Time, and 3% Contract. Highlights an 66% Physical, 3% Hybrid, and 31% Remote job distribution.
AI/LLM Safety Engineer

AI/LLM Safety Engineer

Propio Language Services

Overland Park, KS • On-site, Remote

Full-time

Posted 25 days ago


Job description

Job Type
Full-time
Description
We are seeking an AI/LLM Safety Engineer to join our AI team and take ownership of how safely our models and agents behave in production; with a focus on AI Safety, Trust & Safety, and Responsible AI. You will design the evaluations that catch unsafe behavior, build the guardrails that stop it, and lead the red-teaming that finds the gaps before our users-or attackers-do. Agent safety is the primary focus of this role: you will help ensure that as our systems gain the ability to call tools and take actions, they do so within well-defined, well-tested boundaries.
Key Responsibilities:
LLM Safety Evaluation & Red Teaming
  • Design and maintain a safety evaluation framework-adversarial prompt sets, scenario-based test suites, and regression suites-so that every model and agent update is validated before it ships.
  • Lead structured red-teaming exercises covering jailbreaks, prompt injection, tool misuse, and data exfiltration; document findings and drive each issue through to remediation and closure.

Guardrails & Runtime Controls
  • Build and iterate on guardrail logic, including input/output filtering, tool-boundary constraints, action validation, sensitive-data redaction, and policy prompting.
  • Integrate safety checks into CI/CD and runtime so that unsafe behavior is intercepted before it reaches users.

Agent Safety (primary focus of this role)
  • Perform threat modeling for agentic scenarios: tool-call boundaries, sandbox isolation, and least-privilege access, with particular attention to preventing agents from exfiltrating data or executing irreversible actions through chained tool calls.
  • Conduct safety reviews of reinforcement-learning (RL) environments and trajectory data, partnering with environment and agent engineering teams to embed safety constraints directly into the environments themselves.

Monitoring & Observability
  • Instrument AI features for safety with structured logging, tracing, and metrics, enabling detection of unsafe patterns and regressions in production.

Governance & Collaboration
  • Prepare evidence for governance reviews-test reports, evaluation summaries, and mitigation validation-aligned with internal Responsible AI standards.
  • Collaborate with Product and UX to improve safety interactions (warnings, confirmations, refusal messaging, and feedback collection), and align evaluation goals with the Research and Data teams.

Requirements
  • Bachelor's or Master's degree in Computer Science, Software Engineering, Cybersecurity, or a related technical field-or equivalent practical experience.
  • 4+ years building production software, with direct experience working on-or securing-ML/LLM systems.
  • Strong software engineering skills with the ability to write production-grade code (primarily Python), beyond scripting or notebook prototyping.
  • Solid understanding of LLMs and ML: how models work, prompt engineering, and the safety implications of fine-tuning and RAG (e.g., unsafe retrieval, tool misuse, and data exfiltration).
  • A security mindset with demonstrated threat-modeling ability; able to threat-model AI workflows and familiar with the fundamentals of access control, data retention, and incident response.
  • Familiarity with the LLM attack surface-prompt injection, jailbreaks, data poisoning, and supply-chain risk-and working knowledge of the OWASP LLM Top 10.
  • Hands-on experience with at least one of safety evaluation or red teaming, with the ability to walk through a real finding and how it was remediated.

Preferred Qualifications
  • Hands-on experience with industry safety tooling such as garak, PyRIT, promptfoo, Giskard, and NeMo Guardrails, and the ability to articulate the trade-offs between them.
  • Visible output in AI safety or security: publications at relevant venues (e.g., the NeurIPS AI Safety Workshop, USENIX Security, or DEF CON AI Village), open-source contributions, or responsible disclosures on frontier models with public write-ups.
  • Familiarity with AI governance and compliance frameworks (NIST AI RMF, ISO/IEC 42001, EU AI Act) and the ability to translate compliance requirements into concrete engineering tasks.
  • Engineering experience with agents, RL environments, and/or tool use.
  • Practical experience with threat-modeling methodologies such as MITRE ATLAS and STRIDE/PASTA.

About Propio
Propio is on a mission to make communication accessible to everyone. As a leader in real-time interpretation and multilingual language services, we connect people with the information they need across language, culture, and modality. We are committed to building AI-powered tools that enhance interpreter workflows, automate multilingual insights, and scale communication quality across industries.