LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith ... AI quality and observability: Experience with agent evaluation, prompt regression, model comparison ...
Key responsibilities include continuously improving the design, quality, and reuse of the solution ... Leads the technical oversight for teams in solution development including design reviews and code ...
Key responsibilities include continuously improving the design, quality, and reuse of the solution ... Leads the technical oversight for teams in solution development including design reviews and code ...
Senior ML Ops Engineer
Philadelphia, PA · On-site
$112K - $179K/yr
Build evaluation pipelines: offline IR metrics (NDCG, MAP, MRR), LLM quality metrics (faithfulness ... They publish a vast number of peer-reviewed journals covering a wide range of disciplines. These ...
Senior ML Ops Engineer
Philadelphia, PA · On-site
$112K - $179K/yr
Build evaluation pipelines: offline IR metrics (NDCG, MAP, MRR), LLM quality metrics (faithfulness ... They publish a vast number of peer-reviewed journals covering a wide range of disciplines. These ...
Key responsibilities include continuously improving the design, quality, and reuse of the solution ... Leads the technical oversight for teams in solution development including design reviews and code ...
Key responsibilities include continuously improving the design, quality, and reuse of the solution ... Leads the technical oversight for teams in solution development including design reviews and code ...
Technical Training & Quality Manager
$145K - $175K/yr
LLM evaluation * RLHF / human feedback * Prompt-response evaluation * NLP and text classification ... Participate in client calibration sessions and quality reviews. * Present quality dashboards ...
New
Technical Training & Quality Manager
$145K - $175K/yr
LLM evaluation * RLHF / human feedback * Prompt-response evaluation * NLP and text classification ... Participate in client calibration sessions and quality reviews. * Present quality dashboards ...
New
AI/LLM Safety Engineer
Overland Park, KS · On-site +1
LLM Safety Evaluation & Red Teaming * Design and maintain a safety evaluation framework-adversarial ... quality across industries. Notice of AI Use in Job Application Review As part of our commitment in ...
AI/LLM Safety Engineer
Overland Park, KS · On-site +1
LLM Safety Evaluation & Red Teaming * Design and maintain a safety evaluation framework-adversarial ... quality across industries. Notice of AI Use in Job Application Review As part of our commitment in ...
AI/LLM Safety Engineer
Overland Park, KS · On-site
LLM Safety Evaluation & Red Teaming * Design and maintain a safety evaluation framework ... quality across industries. Notice of AI Use in Job Application Review As part of our commitment in ...
Quick apply
AI/LLM Safety Engineer
Overland Park, KS · On-site
LLM Safety Evaluation & Red Teaming * Design and maintain a safety evaluation framework ... quality across industries. Notice of AI Use in Job Application Review As part of our commitment in ...
Quality & Release Engineering Leader
Atlanta, GA · On-site
$69K - $89K/yr
Job Summary : Bricklayer AI is the first multi-agent LLM-based AI solution that brings autonomous ... quality reviews following significant production issues • Drive continuous improvement ...
Quality & Release Engineering Leader
Atlanta, GA · On-site
$69K - $89K/yr
Job Summary : Bricklayer AI is the first multi-agent LLM-based AI solution that brings autonomous ... quality reviews following significant production issues • Drive continuous improvement ...
Engage with customer technical stakeholders to understand evaluation goals, review methodologies ... Produce high-quality technical documentation, internal research reports, and client-facing ...
Engage with customer technical stakeholders to understand evaluation goals, review methodologies ... Produce high-quality technical documentation, internal research reports, and client-facing ...
LLM Engineer
Houston, TX · On-site
$120K - $130K/yr
We are seeking a detail-oriented LLM Automation Engineer to support AI-driven data analysis ... Ability to review AI outputs for quality assurance and make necessary adjustments * Strong ...
Quick apply
LLM Engineer
Houston, TX · On-site
$120K - $130K/yr
We are seeking a detail-oriented LLM Automation Engineer to support AI-driven data analysis ... Ability to review AI outputs for quality assurance and make necessary adjustments * Strong ...
Participate in system design, code reviews, debugging, and troubleshooting. * Collaborate with ... Ensure code quality through unit testing, integration testing, and best development practices.
Quick apply
Participate in system design, code reviews, debugging, and troubleshooting. * Collaborate with ... Ensure code quality through unit testing, integration testing, and best development practices.
AI/LLM Safety Engineer
Overland Park, KS · On-site
Conduct safety reviews of reinforcement-learning (RL) environments and trajectory data, partnering ... quality across industries.
AI/LLM Safety Engineer
Overland Park, KS · On-site
Conduct safety reviews of reinforcement-learning (RL) environments and trajectory data, partnering ... quality across industries.
AI Quality Analyst
Palo Alto, CA · Remote
$30/hr
LLM Evaluation * Personalization * Prompt Design * Prompt Engineering * Quality Assurance * Remote Contractor * Response Ranking * Side-by-Side Evaluation * SxS Review * Google Account ID:qnkTyx
Quick apply
AI Quality Analyst
Palo Alto, CA · Remote
$30/hr
LLM Evaluation * Personalization * Prompt Design * Prompt Engineering * Quality Assurance * Remote Contractor * Response Ranking * Side-by-Side Evaluation * SxS Review * Google Account ID:qnkTyx
In every relationship, we honor our 36+ year legacy delivering the highest quality data and ... Engage with customer technical stakeholders to understand evaluation goals, review methodologies ...
In every relationship, we honor our 36+ year legacy delivering the highest quality data and ... Engage with customer technical stakeholders to understand evaluation goals, review methodologies ...
Create annotation schemas, labeling guidelines, QA rubrics, golden datasets, and reviewer workflows ... Experience with LLM post-training workflows such as SFT, instruction tuning, preference data, RLHF ...
Create annotation schemas, labeling guidelines, QA rubrics, golden datasets, and reviewer workflows ... Experience with LLM post-training workflows such as SFT, instruction tuning, preference data, RLHF ...
Principal Software Engineer - LLM Optimization
Jersey City, NJ · On-site
$147K - $198K/yr
... code quality, and operational outcomes at scale (e.g., AI-driven PR review assistance, test ... Deep, hands-on experience with LLM inference systems - vLLM, TensorRT-LLM, SGLang, LLM-D or ...
Principal Software Engineer - LLM Optimization
Jersey City, NJ · On-site
$147K - $198K/yr
... code quality, and operational outcomes at scale (e.g., AI-driven PR review assistance, test ... Deep, hands-on experience with LLM inference systems - vLLM, TensorRT-LLM, SGLang, LLM-D or ...
Llm Quality Reviewer information
What is an LLM Quality Reviewer?
What are the key skills and qualifications needed to thrive as an LLM Quality Reviewer, and why are they important?
What are some typical challenges faced by LLM Quality Reviewers, and how can they be addressed?
What is the difference between Llm Quality Reviewer vs Data Annotator?
| Aspect | Llm Quality Reviewer | Data Annotator |
|---|---|---|
| Required Credentials | Basic understanding of AI/ML concepts, attention to detail | Basic computer skills, attention to detail |
| Work Environment | Remote or office-based, focused on reviewing AI outputs | Remote or office-based, focused on labeling data |
| Employer & Industry | Tech companies, AI/ML industry | Tech companies, data labeling services |
| Search & Comparison Intent | Understanding quality review roles in AI | Understanding data labeling roles in AI |
The main difference between an Llm Quality Reviewer and a Data Annotator is that the reviewer assesses and ensures the quality of AI-generated outputs, while the annotator labels and prepares data for training AI models. Both roles require attention to detail and are common in AI/ML industries, but the reviewer focuses on evaluating existing outputs, whereas the annotator creates the training data.
What cities are hiring for Llm Quality Reviewer jobs?
Cities with the most Llm Quality Reviewer job openings:
What states have the most Llm Quality Reviewer jobs?
States with the most job openings for Llm Quality Reviewer jobs include:
What are popular job titles related to Llm Quality Reviewer jobs?
For Llm Quality Reviewer jobs, the most frequently searched job titles are:

Principal Quality Assurance Engineer - Generative AI & LLM Platforms
San Juan, PR • Hybrid
Full-time
Posted 6 days ago
Hewlett Packard Enterprise rating
8.8
Based on 28 frontline employees who took The Breakroom Quiz
Job description
Who We Are:
Hewlett Packard Enterprise is the global edge-to-cloud company advancing the way people live and work. We help companies connect, protect, analyze, and act on their data and applications wherever they live, from edge to cloud, so they can turn insights into outcomes at the speed required to thrive in today's complex world.Our culture thrives onfinding new and better ways to accelerate what's next.We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs.We make bold moves, together, and are a force for good. If you are looking to stretch and grow your career our culture will embrace you.Open up opportunities with HPE.
Job Description:
Job Family Definition:
Designs, develops, troubleshoots and debugs software programs for software enhancements and new products. Develops software including operating systems, compilers, routers, networks, utilities, databases and Internet-related tools. Determines hardware compatibility and/or influences hardware design.
Management Level Definition:
Contributions have visible technical impact on a product or major subcomponent. Applies in-depth professional knowledge and innovative ideas to solve complex problems. Visible contributions improve time-to-market, achieve cost reductions, or satisfy current and future unmet customer needs. Recognized internal authority on key technology area applying innovative principles and ideas. Provides technical leadership for significant project/program work. Leads or participates in cross-functional initiatives and contributes to mentorship and knowledge sharing across the organization.
Role Overview:
We are seeking a Senior Quality Engineer to own quality strategy and hands-on validation for enterprise products built with Generative AI, large language models, distributed systems, and modern React-based user interfaces.
This role combines strong software testing fundamentals with AI evaluation expertise. The successful candidate will design automated quality frameworks, test deterministic and probabilistic behaviour, uncover complex cross-layer defects, influence architecture for testability, and help engineering teams deliver secure, reliable, accessible, high-performance products at speed.
Responsibilities:
Define risk-based test strategies, quality gates, release criteria, traceability, and measurable objectives across AI, UI, API, cloud, and network layers.
Evaluate model quality, groundedness, safety, privacy, robustness, tool use, permissions, failure recovery, latency, and cost.
Validate React interfaces, web standards, accessibility, security, APIs, asynchronous workflows, and enterprise integrations.
Test AWS deployments, distributed systems, traffic behavior, scaling, failover, disaster recovery, and adverse network conditions.
Build reusable Python and pytest frameworks and execute performance, scale, reliability, and end-to-end testing with CI/CD integration.
Use telemetry, incidents, and user feedback to detect regressions, strengthen coverage, and improve preventive controls.
Partner across disciplines, use AI-assisted testing responsibly, mentor engineers, and promote shared ownership of quality.
Education and Experience Required:
Bachelor's or Master's degree in Computer Science, Machine Learning, Artificial Intelligence, or a related discipline.
12+ years' experience with at least 3 years of experience in a lead role
Knowledge and Skills:
Quality Engineering: Enterprise test strategy, automation, quality governance
Automation: Python, pytest, UI/API testing, mocking, diagnostics, CI/CD
Web Testing: React, responsive design, accessibility, cross-browser, web security,
API testing: JMeter, SoapUI, Postman
GenAI Evaluation: LLM testing, groundedness, hallucination, safety, regression, metrics, human review, LangSmith, LangGraph, Langfuse, MCP
Networking & Protocols: IP clos fabric, EVPN, VXLAN, BGP, MPLS, NETCONF, RESTCONF, gRPC, SNMP, LLDP, traffic simulators such as Ixia, Spirent etc...
Performance Engineering: Performance, scale, functional, integration, E2E, regression, security, accessibility, reliability, compatibility
Problem-Solving: Architecture analysis, cross-layer debugging, risk assessment, root-cause communication
AI quality and observability: Experience with agent evaluation, prompt regression, model comparison, tracing, and production monitoring.
Performance and resilience: Proficiency in API and browser performance testing, cloud-scale resilience, and chaos testing.
Security and complex platforms: Knowledge of OWASP risks, threat modelling, adversarial validation, and testing multi-tenant or distributed enterprise architectures.
AWS and monitoring: Familiarity with AI services, container platforms, and observability tools such as Datadog.
Credentials and leadership: Relevant networking certifications and demonstrated leadership across automation, cloud quality, performance, security, or responsible AI.
What We Can Offer You:
Health & Wellbeing
We strive to provide our team members and their loved ones with a comprehensive suite of benefits that supports their physical, financial and emotional wellbeing.
Personal & Professional Development
We also invest in your career because the better you are, the better we all are. We have specific programs catered to helping you reach any career goals you have - whether you want to become a knowledge expert in your field or apply your skills to another division.
Unconditional Inclusion
We are unconditionally inclusive in the way we work and celebrate individual uniqueness. We know varied backgrounds are valued and succeed here. We have the flexibility to manage our work and personal needs. We make bold moves, together, and are a force for good.
Let's Stay Connected:
Follow @HPECareers on Instagram to see the latest on people, culture and tech at HPE.
Job:
EngineeringJob Level:
TCP_05HPE is an Equal Employment Opportunity/ Veterans/Disabled/LGBT employer. We do not discriminate on the basis of race, gender, or any other protected category, and all decisions we make are made on the basis of qualifications, merit, and business need. Our goal is to be one global team that is representative of our customers, in an inclusive environment where we can continue to innovate and grow together. Please click here: Equal Employment Opportunity.
Hewlett Packard Enterprise is EEO Protected Veteran/ Individual with Disabilities.
HPE will comply with all applicable laws related to employer use of arrest and conviction records, including laws requiring employers to consider for employment qualified applicants with criminal histories.
Recruitment Fraud Alert
We have become aware of an increase in fraudulent recruitment activities in which individuals impersonate our company or authorized recruitment agencies to offer fake employment opportunities. These scams may occur through false websites, emails, social media, or chat-based applications and often aim to obtain personal information or money. Please note that Hewlett Packard Enterprise (HPE), its direct and indirect subsidiaries and affiliated companies, and its authorized recruitment agencies/vendors will never charge a candidate a registration fee, hiring fee, or any other fee in connection with its recruitment and hiring process. We also never request personal information such as back account details, Social Security numbers, or national IDs via social media or chat applications.
All legitimate job opportunities will come through official company channels, and candidates are responsible for verifying the credentials of any third party claiming to represent the company. Any reliance on fraudulent communication is at the individual's own risk, and HPE disclaims legal liability for any resulting damages. If you suspect recruitment fraud, do not share personal information or make any payments and report the incident to your local authorities immediately.
What Hewlett Packard Enterprise employees say
Pay
Benefits
Hours and flexibility
Workplace
Get the full story on Breakroom
About Hewlett Packard Enterprise
Sourced by ZipRecruiter