1

Llm Quality Reviewer Jobs (NOW HIRING)

Senior ML Ops Engineer

Philadelphia, PA · On-site

$112K - $179K/yr

Build evaluation pipelines: offline IR metrics (NDCG, MAP, MRR), LLM quality metrics (faithfulness ... They publish a vast number of peer-reviewed journals covering a wide range of disciplines. These ...

LLM evaluation * RLHF / human feedback * Prompt-response evaluation * NLP and text classification ... Participate in client calibration sessions and quality reviews. * Present quality dashboards ...

New

LLM Safety Evaluation & Red Teaming * Design and maintain a safety evaluation framework ... quality across industries. Notice of AI Use in Job Application Review As part of our commitment in ...

LLM Engineer

Houston, TX · On-site

$120K - $130K/yr

We are seeking a detail-oriented LLM Automation Engineer to support AI-driven data analysis ... Ability to review AI outputs for quality assurance and make necessary adjustments * Strong ...

LLM Evaluation * Personalization * Prompt Design * Prompt Engineering * Quality Assurance * Remote Contractor * Response Ranking * Side-by-Side Evaluation * SxS Review * Google Account ID:qnkTyx

Showing results 21-40

Llm Quality Reviewer information

What is an LLM Quality Reviewer?

LLM Quality Reviewers are professionals who evaluate the outputs of large language models (LLMs) to ensure accuracy, relevance, and adherence to guidelines. Their responsibilities include reviewing generated content, detecting biases or errors, and providing feedback to improve model performance. They play a crucial role in maintaining the quality and reliability of AI-generated responses, often collaborating with data scientists and engineers. This position typically requires strong analytical skills, attention to detail, and familiarity with AI technologies.

What are the key skills and qualifications needed to thrive as an LLM Quality Reviewer, and why are they important?

To thrive as an LLM Quality Reviewer, you need a strong background in linguistics, natural language processing, and thorough understanding of large language models, typically supported by a relevant degree or experience in AI or computer science. Familiarity with annotation tools, model evaluation frameworks, and platforms like Python or SQL is important for assessing and improving model outputs. Attention to detail, critical thinking, and clear communication enable effective feedback and collaboration with development teams. These skills ensure accurate evaluation, improved model performance, and alignment with project goals in the fast-evolving AI landscape.

What are some typical challenges faced by LLM Quality Reviewers, and how can they be addressed?

LLM Quality Reviewers often encounter challenges such as evaluating large volumes of AI-generated content under tight deadlines and ensuring consistent application of complex evaluation guidelines. Attention to detail and strong communication skills are essential, as is the ability to provide actionable feedback to model developers. Collaborating closely with data scientists and engineers helps reviewers stay aligned on quality standards and resolve ambiguities promptly. Developing a systematic approach to reviews and staying updated on evolving best practices can make the role more manageable and rewarding.

What is the difference between Llm Quality Reviewer vs Data Annotator?

AspectLlm Quality ReviewerData Annotator
Required CredentialsBasic understanding of AI/ML concepts, attention to detailBasic computer skills, attention to detail
Work EnvironmentRemote or office-based, focused on reviewing AI outputsRemote or office-based, focused on labeling data
Employer & IndustryTech companies, AI/ML industryTech companies, data labeling services
Search & Comparison IntentUnderstanding quality review roles in AIUnderstanding data labeling roles in AI

The main difference between an Llm Quality Reviewer and a Data Annotator is that the reviewer assesses and ensures the quality of AI-generated outputs, while the annotator labels and prepares data for training AI models. Both roles require attention to detail and are common in AI/ML industries, but the reviewer focuses on evaluating existing outputs, whereas the annotator creates the training data.

More about Llm Quality Reviewer jobs

What cities are hiring for Llm Quality Reviewer jobs?

Cities with the most Llm Quality Reviewer job openings:

What states have the most Llm Quality Reviewer jobs?

States with the most job openings for Llm Quality Reviewer jobs include:

What are popular job titles related to Llm Quality Reviewer jobs?

For Llm Quality Reviewer jobs, the most frequently searched job titles are:

Infographic showing various Llm Quality Reviewer job openings in the United States as of September 2026, with employment types broken down into 2% As Needed, 80% Full Time, 14% Part Time, 3% Contract, and 1% Nights. Highlights an 83% Physical, 1% Hybrid, and 16% Remote job distribution.

Principal Quality Assurance Engineer - Generative AI & LLM Platforms

San Juan, PR • Hybrid

Hewlett Packard Enterprise
IT Services • 10K+ employees

Full-time

Posted 6 days ago


Hewlett Packard Enterprise rating

8.8

Company rating: 8.8 out of 10

Based on 28 frontline employees who took The Breakroom Quiz


Job description

Principal Quality Assurance Engineer - Generative AI & LLM PlatformsThis role has been designed as 'Hybrid' with a requirement that you will work on average 2 days per week from an HPE office.

Who We Are:

Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today's complex world.Our culture thrives onfinding new and better ways to accelerate what's next.We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs.We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you.Open up opportunities with HPE.

Job Description:

Job Family Definition:

Designs, develops, troubleshoots and debugs software programs for software enhancements and new products. Develops software including operating systems, compilers, routers, networks, utilities, databases and Internet-related tools. Determines hardware compatibility and/or influences hardware design.

Management Level Definition:

Contributions have visible technical impact on a product or major subcomponent. Applies in-depth professional knowledge and innovative ideas to solve complex problems. Visible contributions improve time-to-market, achieve cost reductions, or satisfy current and future unmet customer needs. Recognized internal authority on key technology area applying innovative principles and ideas. Provides technical leadership for significant project/program work. Leads or participates in cross-functional initiatives and contributes to mentorship and knowledge sharing across the organization.

Role Overview:

We are seeking a Senior Quality Engineer to own quality strategy and hands-on validation for enterprise products built with Generative AI, large language models, distributed systems, and modern React-based user interfaces.

This role combines strong software testing fundamentals with AI evaluation expertise. The successful candidate will design automated quality frameworks, test deterministic and probabilistic behaviour, uncover complex cross-layer defects, influence architecture for testability, and help engineering teams deliver secure, reliable, accessible, high-performance products at speed.

Responsibilities:

  • Define risk-based test strategies, quality gates, release criteria, traceability, and measurable objectives across AI, UI, API, cloud, and network layers.

  • Evaluate model quality, groundedness, safety, privacy, robustness, tool use, permissions, failure recovery, latency, and cost.

  • Validate React interfaces, web standards, accessibility, security, APIs, asynchronous workflows, and enterprise integrations.

  • Test AWS deployments, distributed systems, traffic behavior, scaling, failover, disaster recovery, and adverse network conditions.

  • Build reusable Python and pytest frameworks and execute performance, scale, reliability, and end-to-end testing with CI/CD integration.

  • Use telemetry, incidents, and user feedback to detect regressions, strengthen coverage, and improve preventive controls.

  • Partner across disciplines, use AI-assisted testing responsibly, mentor engineers, and promote shared ownership of quality.

Education and Experience Required:

  • Bachelor's or Master's degree in Computer Science, Machine Learning, Artificial Intelligence, or a related discipline.

  • 12+ years' experience with at least 3 years of experience in a lead role

Knowledge and Skills:

  • Quality Engineering: Enterprise test strategy, automation, quality governance

  • Automation: Python, pytest, UI/API testing, mocking, diagnostics, CI/CD

  • Web Testing: React, responsive design, accessibility, cross-browser, web security,

  • API testing: JMeter, SoapUI, Postman

  • GenAI Evaluation: LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith, LangGraph, Langfuse, MCP

  • Networking & Protocols: IP clos fabric, EVPN, VXLAN, BGP, MPLS, NETCONF, RESTCONF, gRPC, SNMP, LLDP, traffic simulators such as Ixia, Spirent etc...

  • Performance Engineering: Performance, scale, functional, integration, E2E, regression, security, accessibility, reliability, compatibility

  • Problem-Solving: Architecture analysis, cross-layer debugging, risk assessment, root-cause communication

  • AI quality and observability: Experience with agent evaluation, prompt regression, model comparison, tracing, and production monitoring.

  • Performance and resilience: Proficiency in API and browser performance testing, cloud-scale resilience, and chaos testing.

  • Security and complex platforms: Knowledge of OWASP risks, threat modelling, adversarial validation, and testing multi-tenant or distributed enterprise architectures.

  • AWS and monitoring: Familiarity with AI services, container platforms, and observability tools such as Datadog.

  • Credentials and leadership: Relevant networking certifications and demonstrated leadership across automation, cloud quality, performance, security, or responsible AI.

What We Can Offer You:

Health & Wellbeing

We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.

Personal & Professional Development

We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have - whether you want to become a knowledge expert in your field or apply your skills to another division.

Unconditional Inclusion

We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.

Let's Stay Connected:

Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.

Job:

Engineering

Job Level:

TCP_05

HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together. Please click here: Equal Employment Opportunity.

Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.

HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.

Recruitment Fraud Alert

We have become aware of an increase in fraudulent recruitment activities in which individuals impersonate our company or authorized recruitment agencies to offer fake employment opportunities. These scams may occur through false websites, emails, social media, or chat-based applications and often aim to obtain personal information or money. Please note that Hewlett Packard Enterprise (HPE), its direct and indirect subsidiaries and affiliated companies, and its authorized recruitment agencies/vendors will never charge a candidate a registration fee, hiring fee, or any other fee in connection with its recruitment and hiring process. We also never request personal information such as back account details, Social Security numbers, or national IDs via social media or chat applications.

All legitimate job opportunities will come through official company channels, and candidates are responsible for verifying the credentials of any third party claiming to represent the company. Any reliance on fraudulent communication is at the individual's own risk, and HPE disclaims legal liability for any resulting damages. If you suspect recruitment fraud, do not share personal information or make any payments and report the incident to your local authorities immediately.


What Hewlett Packard Enterprise employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom