Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Clinical AI Evaluation Specialist
$90K - $115K/yr
You will design evaluation frameworks, develop and refine automated monitoring approaches (including prompt engineering for evaluation automation), conduct structured reviews of AI-generated content ...
Clinical AI Evaluation Specialist
$90K - $115K/yr
You will design evaluation frameworks, develop and refine automated monitoring approaches (including prompt engineering for evaluation automation), conduct structured reviews of AI-generated content ...
AIML - Software Engineer - AI, Evaluation
Cupertino, CA · On-site
$150K - $277K/yr
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
AIML - Software Engineer - AI, Evaluation
Cupertino, CA · On-site
$150K - $277K/yr
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · Remote
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · Remote
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Quick apply
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Machine Learning Platform Engineer, AI Evaluation Platform (All levels)
Seattle, WA · On-site
$175 - $263.30/hr
Machine Learning Platform Engineer, AI Evaluation Platform (All levels) Seattle, Washington, United States Software and Services Join Apple Services Engineering to build the next generation of AI ...
Machine Learning Platform Engineer, AI Evaluation Platform (All levels)
Seattle, WA · On-site
$175 - $263.30/hr
Machine Learning Platform Engineer, AI Evaluation Platform (All levels) Seattle, Washington, United States Software and Services Join Apple Services Engineering to build the next generation of AI ...
AIML - Software Engineer - AI, Evaluation
$150K - $277K/yr
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
AIML - Software Engineer - AI, Evaluation
$150K - $277K/yr
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · On-site +1
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · On-site +1
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
They are seeking an Eval360 - Error Analysis Engineer to build and operate an evaluation service for AI models, focusing on error analysis and improving model reliability. The role involves ...
They are seeking an Eval360 - Error Analysis Engineer to build and operate an evaluation service for AI models, focusing on error analysis and improving model reliability. The role involves ...
AI Evaluation Subject Matter Expert with Security Clearance
Charleston, SC · On-site
$85K - $117K/yr
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
AI Evaluation Subject Matter Expert with Security Clearance
Charleston, SC · On-site
$85K - $117K/yr
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
Senior Design Evaluation Engineer
Durham, NC · On-site
$90 - $120/hr
ADI combines analog, digital, AI, and software technologies into solutions that combat climate ... Senior Design Evaluation Engineer**The Power Control team is seeking a motivated Design Evaluation ...
Senior Design Evaluation Engineer
Durham, NC · On-site
$90 - $120/hr
ADI combines analog, digital, AI, and software technologies into solutions that combat climate ... Senior Design Evaluation Engineer**The Power Control team is seeking a motivated Design Evaluation ...
Senior Design Evaluation Engineer
$101K - $138K/yr
ADI combines analog, digital, AI, and software technologies into solutions that combat climate ... Senior Design Evaluation Engineer The Power Control team is seeking a motivated Design Evaluation ...
Senior Design Evaluation Engineer
$101K - $138K/yr
ADI combines analog, digital, AI, and software technologies into solutions that combat climate ... Senior Design Evaluation Engineer The Power Control team is seeking a motivated Design Evaluation ...
Senior Design Evaluation Engineer
Durham, NC · On-site
$100K - $135K/yr
ADI combines analog, digital, AI, and software technologies into solutions that combat climate ... Senior Design Evaluation Engineer The Power Control team is seeking a motivated Design Evaluation ...
Senior Design Evaluation Engineer
Durham, NC · On-site
$100K - $135K/yr
ADI combines analog, digital, AI, and software technologies into solutions that combat climate ... Senior Design Evaluation Engineer The Power Control team is seeking a motivated Design Evaluation ...
... Engineer to shape how AI is developed, evaluated, and deployed in healthcare. GUIDE-AI (Guidance ... Establishing evaluation and monitoring methods to assess deployed AI tools at Stanford Health Care ...
... Engineer to shape how AI is developed, evaluated, and deployed in healthcare. GUIDE-AI (Guidance ... Establishing evaluation and monitoring methods to assess deployed AI tools at Stanford Health Care ...
AI Evaluation Scientist with Security Clearance
Fairfax, VA · On-site
$105K - $145K/yr
The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...
AI Evaluation Scientist with Security Clearance
Fairfax, VA · On-site
$105K - $145K/yr
The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...
Fostering cross‐functional collaboration between clinical experts, AI teams, and product developers. Guiding the evolution of evaluation standards for AI‐generated clinical content and medical ...
Fostering cross‐functional collaboration between clinical experts, AI teams, and product developers. Guiding the evolution of evaluation standards for AI‐generated clinical content and medical ...
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are seeking machine learning platform engineers at multiple levels (Mid-Level to Principal) to architect and ...
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are seeking machine learning platform engineers at multiple levels (Mid-Level to Principal) to architect and ...
Research Engineer, GUIDE-AI
Stanford, CA · On-site
$240K/yr
... Engineer to shape how AI is developed, evaluated, and deployed in healthcare. GUIDE-AI (Guidance ... Establishing evaluation and monitoring methods to assess deployed AI tools at Stanford Health Care ...
Research Engineer, GUIDE-AI
Stanford, CA · On-site
$240K/yr
... Engineer to shape how AI is developed, evaluated, and deployed in healthcare. GUIDE-AI (Guidance ... Establishing evaluation and monitoring methods to assess deployed AI tools at Stanford Health Care ...
Fostering cross‐functional collaboration between clinical experts, AI teams, and product developers. Guiding the evolution of evaluation standards for AI‐generated clinical content and medical ...
Fostering cross‐functional collaboration between clinical experts, AI teams, and product developers. Guiding the evolution of evaluation standards for AI‐generated clinical content and medical ...
AI Evaluation Scientist with Security Clearance
Fairfax, VA · On-site
$105K - $145K/yr
The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...
AI Evaluation Scientist with Security Clearance
Fairfax, VA · On-site
$105K - $145K/yr
The AI Evaluation Scientist will work closely with engineers, data scientists, governance analysts, and product teams to develop evaluation metrics, build test harnesses, analyze model behavior, and ...
Ai Evaluation Engineer information
See salary details
$25.48 - $30.14
1% of jobs
$30.14 - $34.79
5% of jobs
$34.79 - $39.44
9% of jobs
$43.46 is the 25th percentile. Wages below this are outliers.
$39.44 - $44.10
12% of jobs
$44.10 - $48.75
10% of jobs
The median wage is $53.08 / hr.
$48.75 - $53.41
15% of jobs
$53.41 - $58.06
15% of jobs
$61.36 is the 75th percentile. Wages above this are outliers.
$58.06 - $62.72
13% of jobs
$62.72 - $67.37
10% of jobs
$67.37 - $72.03
10% of jobs
$72.03 - $76.68
2% of jobs
$25
$53
$76
How much do ai evaluation engineer jobs pay per hour?

Apple rating
8.0
Based on 677 frontline employees who took The Breakroom Quiz
7th of 30 rated technology retailers
Job description
Description
As an AI Software Engineer on the team, you will design and build tools and systems that sit at the intersection of AI modeling, software engineering, and product quality. You will design and develop extensible frameworks, pipelines, and tools that enable efficient development, deployment, and qualitative measurement of AI models. Due to the breadth of products supported, the role requires strong software design and engineering skills. Your work will directly influence product launch decisions and enable teams across Apple to iterate faster and with greater confidence.
Minimum Qualifications
BS/MS/PhD degree in Computer Science, Machine Learning, AI, or a related field.
Exceptional Python skills.
Solid software engineering fundamentals with production experience, including system design, API design, CI/CD, testing strategies, code maintainability, system monitoring, debugging complex systems and etc.
Demonstrated expertise in using AI-assisted software development workflows to accelerate software development while maintaining code quality.
Strong communication skills and proven ability to work collaboratively with cross-functional teams.
Preferred Qualifications
Experience with building LLM applications, frameworks, and offline evaluations.
Familiar with MLOps principles for model lifecycle management.
Experience in building scalable tools for product quality evaluation.
Ability to understand and interpret evaluation reports, including metrics such as precision, recall, run-to-run consistency, and common pitfalls like data leakage.
Product-minded, with a strong ability to translate ambiguous product requirements into solutions.
About Apple
Sourced by ZipRecruiter
Imagine what you could do here! At Apple, new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish. Dynamic, intelligent people and inspiring, innovative technologies are the norm here. The people who work here have reinvented entire industries with all Apple Hardware products. The same real passion for innovation that goes into our products also applies to our practices strengthening our dedication to leave the world better than we found it.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Cupertino, CA, US
Year founded
1976