1

Generative Ai Testing Jobs (NOW HIRING)

As a software engineer, generative ai at WRITER, you'll be at the forefront of expanding human ... Early-detection cancer testing through Galleri * Flexible spending account and dependent FSA ...

... testing, and CI/CD. * Hands-on experience with LLM providers such as Azure OpenAI, OpenAI ... building Generative AI or Machine Learning solutions. * Experience developing enterprise AI ...

Position Description We are seeking a Generative AI Engineer to own the hands-on technical delivery ... Build and maintain automated evaluation pipelines for LLM outputs - prompt regression testing ...

Showing results 41-60

Generative Ai Testing information

See salary details

$31

$53

$76

How much do generative ai testing jobs pay per hour?

As of Sep 10, 2026, the average hourly pay for generative ai testing in the United States is $53.73, according to ZipRecruiter salary data. Most workers in this role earn between $44.23 and $61.54 per hour, depending on experience, location, and employer.

What is generative AI testing?

Generative AI Testing refers to the process of evaluating and validating AI systems, particularly those that generate content such as text, images, or code. This type of testing focuses on assessing the accuracy, reliability, fairness, and safety of generative models to ensure they function as intended and avoid producing harmful or biased outputs. Testers use various methods, including automated and manual techniques, to check for issues like hallucinations, inappropriate content, or security vulnerabilities. The goal is to build trust in generative AI systems and ensure they meet quality and ethical standards before deployment.

What are some common challenges faced when testing generative AI models, and how can I prepare to address them in this role?

Testing generative AI models often involves unique challenges such as evaluating the quality and relevance of generated content, detecting bias or inappropriate outputs, and ensuring model consistency across various prompts. You may work closely with data scientists and engineers to create robust evaluation frameworks and develop automated as well as manual testing strategies. Familiarity with prompt engineering, statistical evaluation techniques, and domain-specific knowledge will help you address these challenges effectively. Proactively staying updated on industry best practices and collaborating with cross-functional teams are key to success in this dynamic field.

What are the key skills and qualifications needed to thrive as a generative AI testing specialist, and why are they important?

To thrive as a Generative AI Testing Specialist, you need a robust understanding of machine learning principles, model evaluation techniques, and a background in computer science or a related field. Familiarity with tools such as Python, TensorFlow, PyTorch, and model evaluation frameworks, as well as experience with automated testing platforms, is typically required. Analytical thinking, attention to detail, and strong communication skills help you identify model weaknesses and collaborate effectively with development teams. These skills are crucial to ensure the reliability, safety, and ethical deployment of generative AI solutions.

What is the difference between Generative Ai Testing vs Data Scientist?

AspectGenerative Ai TestingData Scientist
Required CredentialsKnowledge of AI models, testing tools, programming skillsStatistics, programming, data analysis certifications
Work EnvironmentAI development teams, testing labs, tech companiesResearch labs, tech firms, finance, healthcare
Employer & Industry UsageAI product testing, quality assurance in techData analysis, predictive modeling across industries

Generative Ai Testing focuses on evaluating and validating AI-generated content and models, ensuring quality and accuracy. Data Scientists analyze data, build models, and derive insights. While both roles require programming and AI knowledge, Generative Ai Testing emphasizes testing processes, whereas Data Scientists focus on data analysis and model development.

How do I become a Generative AI Testing?

To become a Generative AI Tester, develop skills in machine learning, natural language processing, and programming languages like Python. Gain experience with AI frameworks such as TensorFlow or PyTorch and understand data quality and model evaluation techniques. Relevant certifications and hands-on projects can enhance your qualifications for roles in AI testing environments.

Is Generative AI Testing a good career?

Generative AI Testing is a growing field within AI development, focusing on evaluating the quality and safety of AI-generated content. It requires skills in machine learning, programming, and understanding AI models, often involving tools like Python and TensorFlow. The role offers opportunities in tech companies and research labs, with demand expected to increase as AI applications expand.
More about Generative Ai Testing jobs

What cities are hiring for Generative Ai Testing jobs?

Cities with the most Generative Ai Testing job openings:

What states have the most Generative Ai Testing jobs?

States with the most job openings for Generative Ai Testing jobs include:

What are popular job titles related to Generative Ai Testing jobs?

For Generative Ai Testing jobs, the most frequently searched job titles are:

Infographic showing various Generative Ai Testing job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 1% As Needed, 83% Full Time, 11% Part Time, 3% Contract, and 1% Nights. Highlights an 90% Physical, 2% Hybrid, and 8% Remote job distribution, with an average salary of $111,750 per year, or $53.7 per hour.

Test Engineer-AI/LLM

Palo Alto, CA • On-site

Full-time

Re-posted 17 days ago


Job description

OPPO US Research Center is seeking a full-time meticulous and innovative AI/LLM Test Engineer to join our cutting-edge AI team. In this critical role, you will evaluate the performance, reliability, and safety of Large Language Models (LLMs) in real-world product scenarios and test end-to-end generative AI solutions. Your work will directly shape how users experience AI-powered features by ensuring robustness, accuracy, and alignment with product goals. This is a unique opportunity to pioneer testing methodologies for next-generation AI systems at the forefront of technology.
We are also seeking a Contractor based LLM Evaluation & QA Engineer to support the testing and validation of large language model (LLM)-powered applications. You will help implement test strategies, execute evaluation workflows, and assist in model performance validation across diverse generative AI use cases.
This contract role is ideal for someone with hands-on experience in AI/ML evaluation, QA engineering, or data analysis who wants to deepen their exposure to generative AI systems.
Requirements
Full-time position requirement:
Core Testing & Evaluation
  • Design and execute performance tests for LLMs across diverse product use cases (e.g., chatbots, content generation etc.).
  • Develop automated test frameworks to evaluate LLM outputs for accuracy, bias, safety, and coherence.
  • Conduct end-to-end testing of integrated generative AI solutions, including APIs, data pipelines, and user interfaces.

Optimization & Validation
  • Collaborate with ML engineers to validate fine-tuned models and optimize prompts for target scenarios.
  • Analyze model failures, edge cases, and adversarial inputs to identify risks and improvement areas.
  • Benchmark LLM performance against industry standards and product-specific KPIs.

Collaboration & Quality Assurance
  • Partner with product, engineering, and research teams to define test requirements and acceptance criteria.
  • Document defects, performance metrics, and test results to drive data-driven improvements.
  • Advocate for AI ethics and safety through rigorous testing of fairness, bias mitigation, and content moderation.

Innovation & Tooling
  • Build scalable tools for synthetic test data generation, prompt variation testing, and automated evaluation workflows.
  • Stay current with advancements in generative AI testing, including red-teaming techniques and evaluation frameworks (e.g., HELM, Dynabench).
  • Propose novel testing strategies for emerging challenges (e.g., hallucinations, context drift).

Basic Qualifications:
  • Bachelor's degree in Computer Science, Data Science, Engineering, or a related technical field, or equivalent practical experience.
  • 1+ years of experience in software testing, data science, or ML validation, with exposure to AI/ML systems.
  • Proficiency in Python and testing frameworks (e.g., PyTest, Selenium).
  • Hands-on experience evaluating LLMs in production environments (e.g., GPT, Claude, Llama, Gemini).
  • Strong analytical skills for dissecting model behavior, statistical performance, and failure modes.
  • Familiarity with cloud platforms (GCP, Azure, or AWS) and MLOps tooling (e.g., MLflow, Weights & Biases).
  • Experience with version control (Git) and agile development methodologies.

Preferred Qualifications:
  • Master's degree in AI, Machine Learning, or a related field.
  • Expertise in prompt engineering, LLM fine-tuning (e.g., LoRA, RLHF), or optimization techniques.
  • Experience with automated evaluation tools (e.g., LangChain, TruLens) or LLM-specific test suites.
  • Knowledge of data pipelines, SQL/NoSQL databases, and API testing (e.g., Postman).
  • Background in statistics, quantitative analysis, or data visualization for test insights.
  • Contributions to AI safety/ethics initiatives or open-source LLM evaluation projects.
  • Experience testing mobile-integrated AI solutions (Android/iOS).

Contractor position requirements:
Testing & Evaluation Support:
  • Execute pre-defined performance tests for LLMs across various tasks (e.g., summarization, Q&A, chatbot flows).
  • Run scripted evaluations to assess outputs for factuality, coherence, and safety.
  • Perform manual and automated test execution on APIs and LLM-integrated user interfaces.

Prompt & model validation:
  • Assist ML engineers in evaluating prompt variations and prompt-tuning outcomes.
  • Log and analyze failure cases, anomalies, and edge cases based on provided guidelines.

Collabration & Documentation
  • Work with QA leads, product managers, and ML engineers to understand test goals and criteria.
  • Report defects, compile evaluation summaries, and maintain testing logs.

Tooling & Antomation:
  • Use existing internal tools or frameworks to automate test runs and result collection.
  • Contribute to prompt generation, input templating, or result tagging processes.

Basic Qualifications:
  • Bachelor's degree or equivalent work experience in a technical field (e.g., Computer Science, Engineering, Data Science).
  • 6+ months experience in software QA, data labeling, LLM evaluation, or ML testing projects.
  • Basic Python proficiency, especially for data processing and automation tasks.
  • Familiarity with LLMs (e.g., GPT, Claude, Gemini) and prompt-based outputs.
  • Comfortable working with tools like Jupyter, Postman, or testing dashboards.
  • Detail-oriented with good documentation habits.

Contractor Details:
  • Duration: Long term
  • Rate: Commensurate with experience
  • Conversion Opportunity: High-performing contractors may be considered for full-time roles

Benefits
OPPO is proud to be an equal opportunity workplace. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements.
The US base salary range for this full-time position is $100,000-$200,000 + bonus + long term incentives benefits. Our salary ranges are determined by role, level, and location.