1

Ai Evaluation Engineer Jobs (NOW HIRING)

... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...

... engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health ... You'll own the evaluation system for the team building the coach: a data scientist partners with ...

Sr. Evaluation Engineer

San Francisco, CA

$123K - $169K/yr

What You'll Do: LogicMonitor is the AI-first hybrid observability platform powering the next ... As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide ...

Sr. Evaluation Engineer

San Francisco, CA · On-site

$123K - $169K/yr

As a Senior AI Engineer, Evaluations, you will design and build the evaluation systems that guide how Edwin AI is developed, tested, and released. You will create production-grade evaluation ...

Showing results 41-60

Ai Evaluation Engineer information

See salary details

$25

$53

$76

How much do ai evaluation engineer jobs pay per hour?

As of Aug 23, 2026, the average hourly pay for ai evaluation engineer in the United States is $53.63, according to ZipRecruiter salary data. Most workers in this role earn between $43.27 and $62.26 per hour, depending on experience, location, and employer.

How to become an AI evaluation engineer?

To become an AI evaluation engineer, candidates typically need a bachelor's or master's degree in computer science, data science, or a related field. Strong skills in machine learning, programming (Python, R), and understanding of AI models are essential, along with experience in data analysis and evaluation metrics. Gaining familiarity with AI frameworks and tools, as well as relevant certifications, can enhance job prospects.

What is the role of AI evaluation engineer?

An AI evaluation engineer is responsible for assessing the performance, accuracy, and fairness of artificial intelligence models. They develop testing protocols, analyze model outputs, and ensure AI systems meet quality and ethical standards, often using tools like benchmarking datasets and evaluation metrics. This role requires strong analytical skills and knowledge of machine learning frameworks.
More about Ai Evaluation Engineer jobs

What cities are hiring for Ai Evaluation Engineer jobs?

Cities with the most Ai Evaluation Engineer job openings:

What states have the most Ai Evaluation Engineer jobs?

States with the most job openings for Ai Evaluation Engineer jobs include:

Infographic showing various Ai Evaluation Engineer job openings in the United States as of August 2026, with employment types broken down into 76% Full Time, 21% Part Time, and 3% Contract. Highlights an 64% Physical, 4% Hybrid, and 32% Remote job distribution, with an average salary of $111,552 per year, or $53.6 per hour.

AI Evaluation Subject Matter Expert

Foxhole Technology

Charleston, SC • Hybrid

Full-time

Re-posted yesterday


Job description

Work Arrangement: Hybrid

Clearance: Active Secret w/TS Capability

Foxhole Technology provides robust cybersecurity and IT support capabilities for federal civilian and defense agencies. A recognized leader in navigating technology and security challenges, Foxhole delivers mission-focused innovations to answer evolving and complex needs. Our talented employee-owners provide agile, scalable services and solutions that solve operational gaps, operate critical systems, and protect and secure the enterprise - across the organization and around the world.

Foxhole Technology is seeking an AI Evaluation SME to join an existing program. The AI Evaluation SME will support the assessment, testing, validation, and operational evaluation of artificial intelligence, machine learning, automation, analytics, and decision-support capabilities being considered for or integrated into the Navy's Next Generation CANES environment. This role will help ensure AI-enabled capabilities are mission-relevant, reliable, secure, explainable, measurable, and suitable for deployment within afloat, tactical, disconnected, intermittent, limited-bandwidth, and multi-security-domain environments.

The SME will develop evaluation frameworks, test methods, metrics, datasets, scenarios, risk assessments, and reporting products that help Navy and CACI stakeholders determine whether AI-enabled capabilities improve network operations, cyber defense, system administration, predictive maintenance, anomaly detection, configuration management, mission planning, or other CANES-related functions.

.

KEY RESPONSIBILITIES: 

  • Serve as a senior technical advisor for AI evaluation, test planning, performance assessment, and operational suitability analysis in support of Next Generation CANES modernization.
  • Develop AI evaluation strategies, test plans, measures of effectiveness, measures of performance, success criteria, risk indicators, and evaluation scorecards.
  • Assess AI, machine learning, generative AI, automation, analytics, and decision-support capabilities for operational relevance, technical maturity, cyber risk, reliability, maintainability, explainability, human oversight, and fleet suitability.
  • Evaluate AI-enabled tools for use cases such as network monitoring, cyber anomaly detection, event correlation, predictive maintenance, help desk automation, configuration compliance, system health monitoring, log analysis, vulnerability prioritization, and operational decision support.
  • Design test scenarios that reflect Navy afloat operating conditions, including limited bandwidth, disconnected operations, contested cyber environments, cross-domain constraints, variable data quality, and platform-specific operational limitations.
  • Define data requirements, ground truth methods, evaluation datasets, labeling approaches, validation methods, and performance baselines for AI-enabled capabilities.
  • Assess AI model performance using appropriate metrics such as accuracy, precision, recall, false positive rate, false negative rate, latency, robustness, drift, confidence calibration, explainability, and operational impact.
  • Evaluate risks associated with hallucination, model brittleness, adversarial manipulation, data poisoning, prompt injection, bias, over-reliance, model drift, cybersecurity exposure, and failure modes in operational environments.
  • Support AI red teaming, cyber survivability assessment, adversarial testing, safety reviews, and responsible AI evaluation activities.
  • Develop human-machine teaming concepts, operator-in-the-loop workflows, trust calibration approaches, escalation procedures, and recommended guardrails for AI-enabled tools.
  • Produce technical reports, evaluation findings, executive summaries, test observations, data analysis products, and recommendations for Navy and CACI leadership.
  • Collaborate with systems engineers, cybersecurity engineers, software developers, data scientists, network engineers, operational testers, fleet users, and government stakeholders.
  • Support technical interchange meetings, design reviews, test readiness reviews, operational assessments, demonstrations, and acquisition decision support.
  • Provide SME input on AI governance, responsible AI implementation, model lifecycle management, configuration control, sustainment, monitoring, and continuous evaluation.

REQUIRED QUALIFICATIONS:

  • Bachelor's degree in computer science, data science, artificial intelligence, engineering, mathematics, statistics, cybersecurity, operations research, information systems, or a related technical discipline preferred. Advanced degree preferred.
  • Additional years of directly relevant AI evaluation, test, cybersecurity, Navy, or DoD mission system experience may be considered in lieu of a degree.
  • Demonstrated experience evaluating AI, machine learning, data analytics, automation, or decision-support systems in defense, intelligence, cybersecurity, network operations, enterprise IT, or mission system environments.
  • Strong understanding of AI / ML evaluation methods, test design, performance metrics, validation approaches, model limitations, and operational risk assessment.
  • Experience developing test plans, evaluation frameworks, measures of effectiveness, measures of performance, data collection plans, and technical reports.
  • Familiarity with cybersecurity, enterprise networks, tactical networks, system monitoring, anomaly detection, log analytics, or network operations use cases.
  • Ability to assess AI-enabled systems in operationally constrained environments, including limited bandwidth, degraded connectivity, edge computing, and mission-critical infrastructure.
  • Understanding of responsible AI concepts, including transparency, explainability, human oversight, robustness, security, bias, accountability, and lifecycle monitoring.
  • Experience working with cross-functional engineering, cyber, data science, software, test, and government stakeholder teams.
  • Strong written and verbal communication skills, including the ability to brief complex AI evaluation findings to technical and non-technical audiences.
  • Active DoD Secret clearance.

DESIRED QUALIFICATIONS: 

  • Experience supporting Navy, DoD, tactical edge, afloat, C4I, cyber, enterprise IT, or mission command systems.
  • Familiarity with CANES, Navy afloat networks, ADNS, NAVWAR programs, RMF, cyber survivability testing, operational test, developmental test, or fleet experimentation.
  • Experience evaluating generative AI, large language models, retrieval-augmented generation, autonomous agents, AI-assisted cyber tools, AI-enabled network operations, or predictive analytics systems.
  • Knowledge of DoD responsible AI guidance, NIST AI Risk Management Framework concepts, RMF, Zero Trust, DevSecOps, MLOps, model monitoring, or secure software supply chain practices.
  • Experience with data analysis tools, scripting, statistical evaluation, dashboards, test automation, or model performance analysis.
  • Experience with AI red teaming, adversarial ML, cyber test events, operational assessments, or acquisition decision support.
  • Top Secret clearance or SCI eligibility.