1

Ai Evaluation Jobs (NOW HIRING)

The Director of AI Evaluation owns how Geisinger defines, proves, and sustains quality across its entire AI portfolio - internally built models and vendor-provided systems alike. This is a hands-on ...

The Director of AI Evaluation owns how Geisinger defines, proves, and sustains quality across its entire AI portfolio - internally built models and vendor-provided systems alike. This is a hands-on ...

Your evaluations and structured feedback will directly contribute to improving AI systems designed to support legal analysis and professional decision-making. This is a fully remote, independent ...

Deposco is seeking an experienced AI Evaluation Engineer to join our innovative quality assurance team. This role focuses on ensuring the accuracy, reliability, and performance of AI-driven ...

Speech AI Evaluation Specialist We are looking for a Speech AI Evaluation Specialist to support the improvement of AI-generated content in Thai or Chinese Simplified. Job Type: Freelance Location:

AI Evaluation Engineer

Washington, DC · On-site

$90 - $120/hr

Ensure enterprise AI platforms and solutions are reliable, tested, and production-ready. Remote / Hybrid Security, Quality & Reliability Full-time Role Overview The QA Engineer is responsible for ...

Assess AI-generated medical responses using structured evaluation frameworks and scoring rubrics. * Identify strengths, weaknesses, omissions, and reasoning errors. * Provide detailed written ...

Finance Domain Expert -- AI Training & Evaluation Type: Contract Compensation: $50-$90/hour Location: Bay Area, California Commitment: 40 hours/week Role Responsibilities * Evaluate the quality of ...

Showing results 21-40

Ai Evaluation information

See salary details

$9

$17

$28

How much do ai evaluation jobs pay per hour?

As of Aug 22, 2026, the average hourly pay for ai evaluation in the United States is $17.85, according to ZipRecruiter salary data. Most workers in this role earn between $16.11 and $18.27 per hour, depending on experience, location, and employer.

What is AI evaluation?

AI evaluation refers to the process of assessing the performance, accuracy, fairness, and reliability of artificial intelligence models or systems. This involves testing AI algorithms using various metrics and datasets to ensure they meet desired standards and function as intended. Evaluators may look for issues like bias, errors, or unintended consequences. The goal is to ensure that AI systems are safe, effective, and trustworthy before they are deployed in real-world applications.

What are the key skills and qualifications needed to thrive as an AI evaluator?

To thrive as an AI Evaluator, you need a strong analytical mindset, attention to detail, and familiarity with AI concepts, often supported by a relevant degree in computer science, linguistics, or a related field. Experience with data annotation platforms, evaluation tools, and sometimes knowledge of programming languages like Python are typically required. Excellent communication skills, critical thinking, and the ability to give constructive feedback are crucial soft skills for this role. These abilities ensure that AI systems are accurately assessed, improved, and aligned with user needs and ethical standards.

What are some common challenges faced by professionals working in AI evaluation roles?

Professionals in AI evaluation often encounter challenges such as ensuring unbiased and accurate assessment of AI models, keeping up with rapidly evolving technologies, and interpreting complex data outputs. Balancing thoroughness with tight project deadlines can be demanding, as can collaborating with cross-functional teams like data scientists, product managers, and engineers to align evaluation metrics with business goals. Additionally, understanding and applying ethical guidelines in AI evaluation is increasingly important as organizations prioritize responsible AI deployment.

What is the difference between Ai Evaluation vs Data Analyst?

AspectAi EvaluationData Analyst
Required CredentialsTypically requires knowledge of AI/ML concepts, certifications in AI tools, and data science fundamentalsRequires degrees in statistics, data analysis, or related fields; often certifications in data analysis tools
Work EnvironmentPrimarily in tech companies, AI development teams, or research labsIn various industries including finance, marketing, healthcare, often in office settings
Employer & Industry UsageUsed by AI development firms, tech companies, and research institutionsUsed across multiple industries for data-driven decision making
Search & Comparison IntentPeople compare to understand AI-specific evaluation rolesOften compared to analyze data trends and insights

While both roles involve working with data, Ai Evaluation focuses on assessing AI models and algorithms, whereas Data Analysts interpret data to inform business decisions. Understanding these differences helps professionals choose the right career path or role based on their skills and industry needs.

Can I get paid to evaluate AI?

Yes, AI evaluation is a recognized job role where professionals assess AI systems for accuracy, bias, and performance. These positions often require skills in data analysis, machine learning, and critical thinking, and may be found in tech companies, research labs, or as freelance opportunities.

How much do AI evaluators make?

AI evaluators typically earn between $40,000 and $80,000 annually, depending on experience, location, and the complexity of tasks. The role often requires skills in data analysis, understanding of AI models, and attention to detail, with some positions offering freelance or part-time options.

How to become an AI evaluation?

To become an AI evaluator, candidates typically need a background in computer science, data analysis, or related fields, along with knowledge of machine learning and AI models. Skills in programming languages like Python, experience with data labeling, and understanding of ethical considerations are also important. Gaining certifications in AI or data science can enhance prospects in this role.

What is an AI evaluationer's job?

An AI evaluationer's job involves assessing the performance, accuracy, and fairness of artificial intelligence models. They analyze outputs, identify biases, and ensure the AI systems meet quality standards, often using specialized tools and datasets. Strong analytical skills and knowledge of machine learning are essential for this role.
More about Ai Evaluation jobs

What cities are hiring for Ai Evaluation jobs?

Cities with the most Ai Evaluation job openings:

What states have the most Ai Evaluation jobs?

States with the most job openings for Ai Evaluation jobs include:

Infographic showing various Ai Evaluation job openings in the United States as of August 2026, with employment types broken down into 76% Full Time, 21% Part Time, and 3% Contract. Highlights an 64% Physical, 4% Hybrid, and 32% Remote job distribution, with an average salary of $37,137 per year, or $17.9 per hour.

Director AI Evaluation

Geisinger Health

Danville, PA • On-site, Remote

Full-time

Medical, Dental, Vision

Re-posted 7 days ago


Geisinger Health rating

6.9

Company rating: 6.9 out of 10

Based on 448 frontline employees who took The Breakroom Quiz

456th of 891 rated healthcare providers


Job description

Location:
Work from home (Pennsylvania)
Shift:
Days (United States of America)
Scheduled Weekly Hours:
40
Worker Type:
Regular
Exemption Status:
Yes
Job Summary:
The Director of AI Evaluation owns how Geisinger defines, proves, and sustains quality across its entire AI portfolio - internally built models and vendor-provided systems alike. This is a hands-on technical leader who also manages the people who do the building and the proving: the data scientists who develop production machine learning and fine-tuned AI systems, and the senior analysts who evaluate them.
Every high-value AI initiative - bought or built - needs a single, credible standard for what constitutes quality, who validates it, and how it stays good in production. The Director sets that standard, leads the team that enforces it, and reports findings to the VP of AI, executive leaders, and the board.
This role is a manager who develops a multidisciplinary team, a technical authority who defines evaluation method across the enterprise, and a quality owner who guides every major AI program toward evidence that withstands scrutiny.
Job Duties:
  • Reports to the VP of AI; directly career-manages the data science line and matrix-manages the Senior Analysts, AI Evaluation.
  • Determines the quality standard for any high-value AI initiative at Geisinger - internally built or vendor-provided - from design through production.
  • Holds bought systems to the same standard as built ones, generating local evidence on whether a tool works here for Geisinger's clinicians and patients rather than accepting vendor aggregate or cherry-picked results.
  • Owns the methodology that holds initiatives to that standard: pre-production validation and live production monitoring.
  • Owns the health of the data science team - attracting and retaining strong technical talent, developing careers, and keeping the bench deep, engaged, and growing - and leads and develops the evaluation team alongside it.
  • Provides hands-on technical guidance to program teams as they design validation studies, equity audits, monitoring plans, and escalation playbooks.
  • Owns the evaluation toolkit and reusable playbooks and templates that let each new program move faster than the last.
  • Translates program-specific failure modes into concrete, measurable production-monitoring metrics; defines what is measured and how, while the AI Platform team builds the backend.
  • Tracks AI System Performance - the single most important accuracy indicator for each system, against thresholds set to clinical tolerance.
  • Tracks User Adoption - engagement, override rates, and time-to-action - distinguishing genuine workflow misalignment and alarm fatigue from poor predictive value.
  • Connects each AI to the Outcome it was deployed to improve (mortality, time-to-treatment, boarding time, denial rate, cost per case) against a pre-launch baseline over a use-case-appropriate horizon, holding both tangible returns and harder-to-quantify value in view.
  • Monitors Equity - the maximum performance gap on the Pillar 1 metric across the subgroups that matter for the initiative, so disparate impact surfaces early.

Work is typically performed in an office environment. Accountable for satisfying all job specific obligations and complying with all organization policies and procedures. The specific statements in this profile are not intended to be all-inclusive. They represent typical elements considered necessary to successfully perform the job.
Position Details:
Required Skills and Qualifications:
  • People-leadership experience - managing, developing, and growing technical staff; building teams, not just leading projects.
  • Strong foundation in experimental design and causal inference, with judgment about which method fits which situation.
  • Hands-on experience designing and running model evaluation studies in real production settings.
  • Experience evaluating LLM or generative AI systems, or comparable experience with complex ML systems where ground truth is ambiguous or noisy.
  • Proven ability to translate ambiguous failure modes into concrete, defensible evaluation designs and monitoring metrics.
  • Strong fluency in Python and SQL; working comfort with modern ML tooling and cloud-native data environments.
  • Experience in evaluating fairness and equity in ML systems.
  • Clear written communication - the role produces evaluation memos and specifications that non-technical decision-makers rely on.
  • Healthcare, clinical, or regulated-industry experience strongly preferred.

Education:
Bachelor's Degree-Related Field of Study (Required)
Experience:
Minimum of 8 years-Related work experience (Required), Minimum of 3 years-Managerial/Supervisory (Required)
Certification(s) and License(s):
Skills:
OUR PURPOSE & VALUES: Everything we do is about caring for our patients, our members, our students, our Geisinger family and our communities.
  • KINDNESS: We strive to treat everyone as we would hope to be treated ourselves.
  • EXCELLENCE: We treasure colleagues who humbly strive for excellence.
  • LEARNING: We share our knowledge with the best and brightest to better prepare the caregivers for tomorrow.
  • INNOVATION: We constantly seek new and better ways to care for our patients, our members, our community, and the nation.
  • SAFETY: We provide a safe environment for our patients and members and the Geisinger family.

We offer healthcare benefits for full time and part time positions from day one, including vision, dental and domestic partners. Perhaps just as important, we encourage an atmosphere of collaboration, cooperation and collegiality.
We know that a diverse workforce with unique experiences and backgrounds makes our team stronger. Our patients, members and community come from a wide variety of backgrounds, and it takes a diverse workforce to make better health easier for all. We are proud to be an affirmative action, equal opportunity employer and all qualified applicants will receive consideration for employment regardless to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or status as a protected veteran.

What Geisinger Health employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom