The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Queens, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Queens, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
New York, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
New York, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Queens, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Queens, NY · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Houston, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Houston, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Austin, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Austin, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Houston, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
AI Evaluation Engineer (Python, QA or Security)
Houston, TX · Remote
$50/hr
What this opportunity involves We're building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks. You'll create challenging tasks and evaluation criteria ...
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation ... Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists ...
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$150 - $230/hr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$150 - $230/hr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$147K - $232K/yr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Software Engineer, AI Evaluation
San Francisco, CA · On-site
$147K - $232K/yr
... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are building the scientific foundation and self-service tools for how AI evaluation is done at scale ...
Join Apple Services Engineering to build the next generation of AI evaluation systems. We are building the scientific foundation and self-service tools for how AI evaluation is done at scale ...
Sr. Evaluation Engineer
$123K - $169K/yr
What You'll Do: LogicMonitor is the AI-first hybrid observability platform powering the next ... As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide ...
Sr. Evaluation Engineer
$123K - $169K/yr
What You'll Do: LogicMonitor is the AI-first hybrid observability platform powering the next ... As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide ...
AIML - Software Engineer - AI, Evaluation
$150K - $277K/yr
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
AIML - Software Engineer - AI, Evaluation
$150K - $277K/yr
Our team specializes in building LLM-as-judge and related tools to improve both the quality and efficiency of these evaluations. We are seeking a highly innovative and passionate AI software engineer ...
Sr. Evaluation Engineer
San Francisco, CA · On-site
$123K - $169K/yr
As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation ...
Sr. Evaluation Engineer
San Francisco, CA · On-site
$123K - $169K/yr
As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · Remote
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · Remote
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Quick apply
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · On-site +1
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Senior Python Developer (AI Evaluation & Benchmarking)
Texas City, TX · On-site +1
$80 - $100/hr
The selected consultants will contribute to AI research by designing programming benchmarks, evaluating AI-generated code, and helping improve the performance, reasoning, and reliability of frontier ...
Machine Learning Platform Engineer, AI Evaluation Platform (All levels)
Seattle, WA · On-site
$175 - $263.30/hr
Machine Learning Platform Engineer, AI Evaluation Platform (All levels) Seattle, Washington, United States Software and Services Join Apple Services Engineering to build the next generation of AI ...
Machine Learning Platform Engineer, AI Evaluation Platform (All levels)
Seattle, WA · On-site
$175 - $263.30/hr
Machine Learning Platform Engineer, AI Evaluation Platform (All levels) Seattle, Washington, United States Software and Services Join Apple Services Engineering to build the next generation of AI ...
Ai Evaluation Engineer information
See salary details
$25.48 - $30.14
1% of jobs
$30.14 - $34.79
5% of jobs
$34.79 - $39.44
9% of jobs
$43.46 is the 25th percentile. Wages below this are outliers.
$39.44 - $44.10
12% of jobs
$44.10 - $48.75
10% of jobs
The median wage is $53.08 / hr.
$48.75 - $53.41
15% of jobs
$53.41 - $58.06
15% of jobs
$61.36 is the 75th percentile. Wages above this are outliers.
$58.06 - $62.72
13% of jobs
$62.72 - $67.37
10% of jobs
$67.37 - $72.03
10% of jobs
$72.03 - $76.68
2% of jobs
$25
$53
$76
How much do ai evaluation engineer jobs pay per hour?
How to become an AI evaluation engineer?
What is the role of AI evaluation engineer?
What cities are hiring for Ai Evaluation Engineer jobs?
Cities with the most Ai Evaluation Engineer job openings:
What states have the most Ai Evaluation Engineer jobs?
States with the most job openings for Ai Evaluation Engineer jobs include:
What job categories do people searching Ai Evaluation Engineer jobs look for?
The top searched job categories for Ai Evaluation Engineer jobs are:

Full-time
Re-posted yesterday
Job description
Work Arrangement: Hybrid
Clearance: Active Secret w/TS Capability
Foxhole Technology provides robust cybersecurity and IT support capabilities for federal civilian and defense agencies. A recognized leader in navigating technology and security challenges, Foxhole delivers mission-focused innovations to answer evolving and complex needs. Our talented employee-owners provide agile, scalable services and solutions that solve operational gaps, operate critical systems, and protect and secure the enterprise - across the organization and around the world.
Foxhole Technology is seeking an AI Evaluation SME to join an existing program. The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation of artificial intelligence, machine learning, automation, analytics, and decision-support capabilities being considered for or integrated into the Navy's Next Generation CANES environment. This role will help ensure AI-enabled capabilities are mission-relevant, reliable, secure, explainable, measurable, and suitable for deployment within afloat, tactical, disconnected, intermittent, limited-bandwidth, and multi-security-domain environments.
The SME will develop evaluation frameworks, test methods, metrics, datasets, scenarios, risk assessments, and reporting products that help Navy and CACI stakeholders determine whether AI-enabled capabilities improve network operations, cyber defense, system administration, predictive maintenance, anomaly detection, configuration management, mission planning, or other CANES-related functions.
.
KEY RESPONSIBILITIES:
- Serve as a senior technical advisor for AI evaluation, test planning, performance assessment, and operational suitability analysis in support of Next Generation CANES modernization.
- Develop AI evaluation strategies, test plans, measures of effectiveness, measures of performance, success criteria, risk indicators, and evaluation scorecards.
- Assess AI, machine learning, generative AI, automation, analytics, and decision-support capabilities for operational relevance, technical maturity, cyber risk, reliability, maintainability, explainability, human oversight, and fleet suitability.
- Evaluate AI-enabled tools for use cases such as network monitoring, cyber anomaly detection, event correlation, predictive maintenance, help desk automation, configuration compliance, system health monitoring, log analysis, vulnerability prioritization, and operational decision support.
- Design test scenarios that reflect Navy afloat operating conditions, including limited bandwidth, disconnected operations, contested cyber environments, cross-domain constraints, variable data quality, and platform-specific operational limitations.
- Define data requirements, ground truth methods, evaluation datasets, labeling approaches, validation methods, and performance baselines for AI-enabled capabilities.
- Assess AI model performance using appropriate metrics such as accuracy, precision, recall, false positive rate, false negative rate, latency, robustness, drift, confidence calibration, explainability, and operational impact.
- Evaluate risks associated with hallucination, model brittleness, adversarial manipulation, data poisoning, prompt injection, bias, over-reliance, model drift, cybersecurity exposure, and failure modes in operational environments.
- Support AI red teaming, cyber survivability assessment, adversarial testing, safety reviews, and responsible AI evaluation activities.
- Develop human-machine teaming concepts, operator-in-the-loop workflows, trust calibration approaches, escalation procedures, and recommended guardrails for AI-enabled tools.
- Produce technical reports, evaluation findings, executive summaries, test observations, data analysis products, and recommendations for Navy and CACI leadership.
- Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists, network engineers, operational testers, fleet users, and government stakeholders.
- Support technical interchange meetings, design reviews, test readiness reviews, operational assessments, demonstrations, and acquisition decision support.
- Provide SME input on AI governance, responsible AI implementation, model lifecycle management, configuration control, sustainment, monitoring, and continuous evaluation.
REQUIRED QUALIFICATIONS:
- Bachelor's degree in computer science, data science, artificial intelligence, engineering, mathematics, statistics, cybersecurity, operations research, information systems, or a related technical discipline preferred. Advanced degree preferred.
- Additional years of directly relevant AI evaluation, test, cybersecurity, Navy, or DoD mission system experience may be considered in lieu of a degree.
- Demonstrated experience evaluating AI, machine learning, data analytics, automation, or decision-support systems in defense, intelligence, cybersecurity, network operations, enterprise IT, or mission system environments.
- Strong understanding of AI / ML evaluation methods, test design, performance metrics, validation approaches, model limitations, and operational risk assessment.
- Experience developing test plans, evaluation frameworks, measures of effectiveness, measures of performance, data collection plans, and technical reports.
- Familiarity with cybersecurity, enterprise networks, tactical networks, system monitoring, anomaly detection, log analytics, or network operations use cases.
- Ability to assess AI-enabled systems in operationally constrained environments, including limited bandwidth, degraded connectivity, edge computing, and mission-critical infrastructure.
- Understanding of responsible AI concepts, including transparency, explainability, human oversight, robustness, security, bias, accountability, and lifecycle monitoring.
- Experience working with cross-functional engineering, cyber, data science, software, test, and government stakeholder teams.
- Strong written and verbal communication skills, including the ability to brief complex AI evaluation findings to technical and non-technical audiences.
- Active DoD Secret clearance.
DESIRED QUALIFICATIONS:
- Experience supporting Navy, DoD, tactical edge, afloat, C4I, cyber, enterprise IT, or mission command systems.
- Familiarity with CANES, Navy afloat networks, ADNS, NAVWAR programs, RMF, cyber survivability testing, operational test, developmental test, or fleet experimentation.
- Experience evaluating generative AI, large language models, retrieval-augmented generation, autonomous agents, AI-assisted cyber tools, AI-enabled network operations, or predictive analytics systems.
- Knowledge of DoD responsible AI guidance, NIST AI Risk Management Framework concepts, RMF, Zero Trust, DevSecOps, MLOps, model monitoring, or secure software supply chain practices.
- Experience with data analysis tools, scripting, statistical evaluation, dashboards, test automation, or model performance analysis.
- Experience with AI red teaming, adversarial ML, cyber test events, operational assessments, or acquisition decision support.
- Top Secret clearance or SCI eligibility.
About Foxhole Technology
Sourced by ZipRecruiter
Company size
51 - 200 Employees
Headquarters location
Fairfax, VA, US
Year founded
2007