The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
AI Evaluations Engineer
San Francisco, CA · On-site
$160K - $210K/yr
# AI Evaluations Engineer## LocationSan Francisco - Hybrid## Employment TypeFull time## Location ... You'll also build and maintain evaluation frameworks for individual system components such as ASR ...
AI Evaluations Engineer
San Francisco, CA · On-site
$160K - $210K/yr
# AI Evaluations Engineer## LocationSan Francisco - Hybrid## Employment TypeFull time## Location ... You'll also build and maintain evaluation frameworks for individual system components such as ASR ...
Director- Applied AI
Lisle, IL · Hybrid
Hold the AI Evaluation Engineer accountable for a durable eval harness -- offline golden datasets, regression suites, LLM-as-judge scoring, human-review queues, and online quality monitoring. * Treat ...
Quick apply
Director- Applied AI
Lisle, IL · Hybrid
Hold the AI Evaluation Engineer accountable for a durable eval harness -- offline golden datasets, regression suites, LLM-as-judge scoring, human-review queues, and online quality monitoring. * Treat ...
$160K - $210K/yr
# AI Evaluations Engineer## LocationSan Francisco - Hybrid## Employment TypeFull time## Location ... You'll also build and maintain evaluation frameworks for individual system components such as ASR ...
$160K - $210K/yr
# AI Evaluations Engineer## LocationSan Francisco - Hybrid## Employment TypeFull time## Location ... You'll also build and maintain evaluation frameworks for individual system components such as ASR ...
Director- Applied AI
Lisle, IL · On-site
$250K/yr
Hold the AI Evaluation Engineer accountable for a durable eval harness - offline golden datasets, regression suites, LLM-as-judge scoring, human-review queues, and online quality monitoring. * Treat ...
Director- Applied AI
Lisle, IL · On-site
$250K/yr
Hold the AI Evaluation Engineer accountable for a durable eval harness - offline golden datasets, regression suites, LLM-as-judge scoring, human-review queues, and online quality monitoring. * Treat ...
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Director- Applied AI
Lisle, IL · Hybrid
$250K/yr
Hold the AI Evaluation Engineer accountable for a durable eval harness - offline golden datasets, regression suites, LLM-as-judge scoring, human-review queues, and online quality monitoring. * Treat ...
Director- Applied AI
Lisle, IL · Hybrid
$250K/yr
Hold the AI Evaluation Engineer accountable for a durable eval harness - offline golden datasets, regression suites, LLM-as-judge scoring, human-review queues, and online quality monitoring. * Treat ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$147K - $232K/yr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$147K - $232K/yr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences.
The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences.
The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences.
New
The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences.
New
Lead quality engineering systems that directly improve AI output reliability, evaluation accuracy, and benchmark performance for frontier lab partners * Build and scale an organization across ...
Lead quality engineering systems that directly improve AI output reliability, evaluation accuracy, and benchmark performance for frontier lab partners * Build and scale an organization across ...
$150K - $277K/yr
AIML - Software Engineer - AI, Evaluation Cupertino, California, United States Machine Learning and AI Do you get excited by building software systems to enhance the automatic evaluation of various ...
$150K - $277K/yr
AIML - Software Engineer - AI, Evaluation Cupertino, California, United States Machine Learning and AI Do you get excited by building software systems to enhance the automatic evaluation of various ...
Engineering Manager, AI Evaluation
San Francisco, CA · On-site
$240K - $280K/yr
Background at companies where AI is the core product; you've shipped in environments where model quality, data pipelines, and evaluation rigor are table stakes * Strong product engineering instincts ...
Quick apply
Engineering Manager, AI Evaluation
San Francisco, CA · On-site
$240K - $280K/yr
Background at companies where AI is the core product; you've shipped in environments where model quality, data pipelines, and evaluation rigor are table stakes * Strong product engineering instincts ...
Lead Forward Deployed Engineer, AI Evaluation Platform
Seattle, WA · On-site
$175K - $263K/yr
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are building the scientific foundation and self-service tools for how AI evaluation is done at scale ...
Lead Forward Deployed Engineer, AI Evaluation Platform
Seattle, WA · On-site
$175K - $263K/yr
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are building the scientific foundation and self-service tools for how AI evaluation is done at scale ...
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
AIML - Software Engineer - AI, Evaluation
Cupertino, CA · On-site
$150K - $277K/yr
AIML - Software Engineer - AI, Evaluation Cupertino, California, United States Machine Learning and AI Do you get excited by building software systems to enhance the automatic evaluation of various ...
AIML - Software Engineer - AI, Evaluation
Cupertino, CA · On-site
$150K - $277K/yr
AIML - Software Engineer - AI, Evaluation Cupertino, California, United States Machine Learning and AI Do you get excited by building software systems to enhance the automatic evaluation of various ...
Partner with AI Engineering, Risk, Compliance, Legal, and Model Governance teams to translate regulatory and business requirements into repeatable AI testing methodologies, vendor evaluation ...
Partner with AI Engineering, Risk, Compliance, Legal, and Model Governance teams to translate regulatory and business requirements into repeatable AI testing methodologies, vendor evaluation ...
Sr. Evaluation Engineer
San Francisco, CA · On-site
$123K - $169K/yr
What You'll Do: LogicMonitor is the AI-first hybrid observability platform powering the next ... As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide ...
Sr. Evaluation Engineer
San Francisco, CA · On-site
$123K - $169K/yr
What You'll Do: LogicMonitor is the AI-first hybrid observability platform powering the next ... As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · Remote
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · Remote
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Quick apply
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Ai Evaluation Engineer information
See salary details
$25.48 - $30.14
1% of jobs
$30.14 - $34.79
5% of jobs
$34.79 - $39.44
9% of jobs
$43.46 is the 25th percentile. Wages below this are outliers.
$39.44 - $44.10
12% of jobs
$44.10 - $48.75
10% of jobs
The median wage is $53.08 / hr.
$48.75 - $53.41
15% of jobs
$53.41 - $58.06
15% of jobs
$61.36 is the 75th percentile. Wages above this are outliers.
$58.06 - $62.72
13% of jobs
$62.72 - $67.37
10% of jobs
$67.37 - $72.03
10% of jobs
$72.03 - $76.68
2% of jobs
$25
$53
$76
How much do ai evaluation engineer jobs pay per hour?
How to become an AI evaluation engineer?
What is the role of AI evaluation engineer?
What cities are hiring for Ai Evaluation Engineer jobs?
Cities with the most Ai Evaluation Engineer job openings:
What states have the most Ai Evaluation Engineer jobs?
States with the most job openings for Ai Evaluation Engineer jobs include:
What are popular job titles related to Ai Evaluation Engineer jobs?
For Ai Evaluation Engineer jobs, the most frequently searched job titles are:

AI Evaluation Subject Matter Expert
Charleston, SC • On-site
Full-time
Re-posted 23 days ago
Key responsibilities
Support the assessment, testing, validation, and operational evaluation of AI, machine learning, automation, analytics, and decision-support capabilities for the Navy's Next Generation CANES environment.
Develop evaluation frameworks, test methods, metrics, datasets, scenarios, risk assessments, and reporting products to determine AI-enabled capabilities' effectiveness, security, and suitability.
Evaluate AI tools and models for operational relevance, technical maturity, cyber risk, reliability, explainability, and fleet suitability, including designing test scenarios reflecting Navy afloat conditions.
Job description
Clearance: Active Secret w/TS Capability
Foxhole Technology provides robust cybersecurity and IT support capabilities for federal civilian and defense agencies. A recognized leader in navigating technology and security challenges, Foxhole delivers mission-focused innovations to answer evolving and complex needs. Our talented employee-owners provide agile, scalable services and solutions that solve operational gaps, operate critical systems, and protect and secure the enterprise - across the organization and around the world
Foxhole Technology is seeking an AI Evaluation SME to join an existing program. The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation of artificial intelligence, machine learning, automation, analytics, and decision-support capabilities being considered for or integrated into the Navy's Next Generation CANES environment. This role will help ensure AI-enabled capabilities are mission-relevant, reliable, secure, explainable, measurable, and suitable for deployment within afloat, tactical, disconnected, intermittent, limited-bandwidth, and multi-security-domain environments.
The SME will develop evaluation frameworks, test methods, metrics, datasets, scenarios, risk assessments, and reporting products that help Navy and CACI stakeholders determine whether AI-enabled capabilities improve network operations, cyber defense, system administration, predictive maintenance, anomaly detection, configuration management, mission planning, or other CANES-related functions.
KEY RESPONSIBILITIES:
- Serve as a senior technical advisor for AI evaluation, test planning, performance assessment, and operational suitability analysis in support of Next Generation CANES modernization.
- Develop AI evaluation strategies, test plans, measures of effectiveness, measures of performance, success criteria, risk indicators, and evaluation scorecards.
- Assess AI, machine learning, generative AI, automation, analytics, and decision-support capabilities for operational relevance, technical maturity, cyber risk, reliability, maintainability, explainability, human oversight, and fleet suitability.
- Evaluate AI-enabled tools for use cases such as network monitoring, cyber anomaly detection, event correlation, predictive maintenance, help desk automation, configuration compliance, system health monitoring, log analysis, vulnerability prioritization, and operational decision support.
- Design test scenarios that reflect Navy afloat operating conditions, including limited bandwidth, disconnected operations, contested cyber environments, cross-domain constraints, variable data quality, and platform-specific operational limitations.
- Define data requirements, ground truth methods, evaluation datasets, labeling approaches, validation methods, and performance baselines for AI-enabled capabilities.
- Assess AI model performance using appropriate metrics such as accuracy, precision, recall, false positive rate, false negative rate, latency, robustness, drift, confidence calibration, explainability, and operational impact.
- Evaluate risks associated with hallucination, model brittleness, adversarial manipulation, data poisoning, prompt injection, bias, over-reliance, model drift, cybersecurity exposure, and failure modes in operational environments.
- Support AI red teaming, cyber survivability assessment, adversarial testing, safety reviews, and responsible AI evaluation activities.
- Develop human-machine teaming concepts, operator-in-the-loop workflows, trust calibration approaches, escalation procedures, and recommended guardrails for AI-enabled tools.
- Produce technical reports, evaluation findings, executive summaries, test observations, data analysis products, and recommendations for Navy and CACI leadership.
- Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists, network engineers, operational testers, fleet users, and government stakeholders.
- Support technical interchange meetings, design reviews, test readiness reviews, operational assessments, demonstrations, and acquisition decision support.
- Provide SME input on AI governance, responsible AI implementation, model lifecycle management, configuration control, sustainment, monitoring, and continuous evaluation.
REQUIRED QUALIFICATIONS:
- Bachelor's degree in computer science, data science, artificial intelligence, engineering, mathematics, statistics, cybersecurity, operations research, information systems, or a related technical discipline preferred. Advanced degree preferred.
- Additional years of directly relevant AI evaluation, test, cybersecurity, Navy, or DoD mission system experience may be considered in lieu of a degree.
- Demonstrated experience evaluating AI, machine learning, data analytics, automation, or decision-support systems in defense, intelligence, cybersecurity, network operations, enterprise IT, or mission system environments.
- Strong understanding of AI / ML evaluation methods, test design, performance metrics, validation approaches, model limitations, and operational risk assessment.
- Experience developing test plans, evaluation frameworks, measures of effectiveness, measures of performance, data collection plans, and technical reports.
- Familiarity with cybersecurity, enterprise networks, tactical networks, system monitoring, anomaly detection, log analytics, or network operations use cases.
- Ability to assess AI-enabled systems in operationally constrained environments, including limited bandwidth, degraded connectivity, edge computing, and mission-critical infrastructure.
- Understanding of responsible AI concepts, including transparency, explainability, human oversight, robustness, security, bias, accountability, and lifecycle monitoring.
- Experience working with cross-functional engineering, cyber, data science, software, test, and government stakeholder teams.
- Strong written and verbal communication skills, including the ability to brief complex AI evaluation findings to technical and non-technical audiences.
- Active DoD Secret clearance.
DESIRED QUALIFICATIONS:
- Experience supporting Navy, DoD, tactical edge, afloat, C4I, cyber, enterprise IT, or mission command systems.
- Familiarity with CANES, Navy afloat networks, ADNS, NAVWAR programs, RMF, cyber survivability testing, operational test, developmental test, or fleet experimentation.
- Experience evaluating generative AI, large language models, retrieval-augmented generation, autonomous agents, AI-assisted cyber tools, AI-enabled network operations, or predictive analytics systems.
- Knowledge of DoD responsible AI guidance, NIST AI Risk Management Framework concepts, RMF, Zero Trust, DevSecOps, MLOps, model monitoring, or secure software supply chain practices.
- Experience with data analysis tools, scripting, statistical evaluation, dashboards, test automation, or model performance analysis.
- Experience with AI red teaming, adversarial ML, cyber test events, operational assessments, or acquisition decision support.
- Top Secret clearance or SCI eligibility.
Requirements of position: Think analytically, effective verbal and written communication skills, make decisions, observe/remember details, interpret data, concentrate on tasks, adjust to change, handle stress/emotions. Regular attendance, maintain work schedule, attend meetings, meet deadlines, keyboard/type, handle confidential information, use math/calculations, stay organized, operate office equipment, may direct others. May be exposed to dust/dirt, humidity, and noise
Foxhole Technology is an Equal Opportunity Employer and makes hiring decisions without regard to race, color, religion, sex (including pregnancy, childbirth and sexual orientation), national origin, age, disability, genetic information, military/veteran status, or any other protected class.
About Foxhole Technology
Sourced by ZipRecruiter
Company size
51 - 200 Employees
Headquarters location
Fairfax, VA, US
Year founded
2007