1

Senior Ai Validation Jobs (NOW HIRING)

About You In order to set you up for success as a Systems Engineer (Senior / Staff), AI Validation at Wayve, we're looking for the following skills and experience. Essential * 5+ years of experience ...

About You In order to set you up for success as a Systems Engineer (Senior / Staff), AI Validation at Wayve, we're looking for the following skills and experience. Essential * 5+ years of experience ...

Senior AI Engineer

Washington, DC · On-site

$118K - $162K/yr

Senior AI Engineer Hybrid - Washington D.C. (preferred) or New York, NY About The Role We are ... Solid understanding of model validation, performance monitoring, and explainability techniques

Senior AI Engineer

Washington, DC · On-site

$118K - $162K/yr

Senior AI Engineer Hybrid - Washington D.C. (preferred) or New York, NY About The Role We are ... Solid understanding of model validation, performance monitoring, and explainability techniques

Senior AI Engineer

Washington, DC · On-site

$118K - $162K/yr

Senior AI Engineer Hybrid - Washington D.C. (preferred) or New York, NY About The Role We are ... Solid understanding of model validation, performance monitoring, and explainability techniques

Sr AI Engineer

Austin, TX · On-site

$122K - $161K/yr

Sr AI Engineer Location: Onsite - San Francisco, CA / Dallas, TX / Austin, TX We are looking for a ... Review and validate AI-generated code for quality, security and performance * Generate and improve ...

next page

Showing results 1-20

Senior Ai Validation information

See salary details

$34

$64

$98

How much do senior ai validation jobs pay per hour?

As of Aug 4, 2026, the average hourly pay for senior ai validation in the United States is $64.77, according to ZipRecruiter salary data. Most workers in this role earn between $51.68 and $73.56 per hour, depending on experience, location, and employer.
More about Senior Ai Validation jobs
What cities are hiring for Senior Ai Validation jobs? Cities with the most Senior Ai Validation job openings:
What are the most commonly searched types of Ai Validation jobs? The most popular types of Ai Validation jobs are:
What states have the most Senior Ai Validation jobs? States with the most job openings for Senior Ai Validation jobs include:
Infographic showing various Senior Ai Validation job openings in the United States as of July 2026, with employment types broken down into 67% Full Time, 22% Part Time, and 11% Contract. Highlights an 59% Physical, 3% Hybrid, and 38% Remote job distribution, with an average salary of $134,726 per year, or $64.8 per hour.

Senior Manager, Engineering - AI Validation

Sonatus

Sunnyvale, CA

Other

Medical, Dental, Vision, Life, Retirement, PTO

Posted 17 days ago


Job description

The Opportunity:

Sonatus is looking for an experienced Senior Engineering Manager to build and lead our AI Validation function - the team responsible for how we test, evaluate, and govern the AI models and agentic capabilities embedded in our software-defined vehicle and cloud platforms, as well as our cloud-only AI and LLM-based products. You will deeply understand how AI is developed and deployed across Sonatus's embedded, cloud, and LLM/RAG-driven environments, identify gaps and friction in current validation practices, and turn those insights into scalable test strategy, evaluation frameworks, and governance mechanisms that let Sonatus ship trustworthy AI-driven features at automotive scale. You will lead and grow a high-performing AI validation team and partner closely with engineering, product, and safety stakeholders to make AI quality and safety a competitive advantage.

Role and Responsibility:
  • Define and drive Sonatus's AI validation strategy across embedded, in-vehicle, cloud-connected, and cloud-native AI systems, identifying gaps in model development, testing, deployment, and governance.
  • Lead, hire, mentor, and grow a high-performing AI Validation organization, establishing scalable engineering processes, technical direction, and execution excellence.
  • Own the end-to-end validation strategy for AI/ML models, LLMs, RAG pipelines, and agentic AI workflows-from data pipelines and model training through cloud services and in-vehicle deployment.
  • Architect and operationalize scalable evaluation frameworks and benchmarking platforms for AI systems, including multi-step agentic workflows, using deterministic metrics, LLM-as-a-Judge methodologies, automated regression testing, and production feedback loops.
  • Design and maintain evaluation harnesses for RAG and agentic systems that measure retrieval quality, grounding, citation accuracy, factual consistency, context relevance, safety, latency, reliability, and execution correctness.
  • Evaluate, integrate, and optimize open-source and commercial AI validation technologies, driving build-versus-buy decisions for Sonatus's AI quality platform.
  • Establish AI governance, Responsible AI practices, model lineage, safety guardrails, and compliance processes appropriate for automotive safety-critical systems and enterprise AI products.
  • Act as the quality gatekeeper for AI-enabled releases, partnering with engineering, product, safety, and OEM stakeholders to identify risks, define release criteria, and ensure production readiness.
  • Collaborate across engineering teams to define validation strategies for emerging AI capabilities, rapidly prototype new evaluation approaches, and standardize successful practices into reusable frameworks.
  • Drive continuous improvement by tracking industry advances in AI evaluation, agentic AI, LLM validation, and RAG systems, translating them into scalable validation capabilities across Sonatus.
Qualifications:
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field required (MS preferred).
  • 10+ years of experience in software or systems engineering-including embedded, cloud, networking, security, or automotive domains-with 3+ years leading high-performing engineering or QA organizations.
  • Hands-on experience developing, deploying, testing, or operating AI/ML systems, with strong expertise in modern ML workflows, neural networks, and MLOps.
  • Deep understanding of LLMs, RAG architectures, vector databases, embeddings, retrieval optimization, and agentic AI frameworks such as LangGraph or equivalent orchestration platforms.
  • Proven experience designing and implementing scalable evaluation frameworks for AI systems, including multi-step agentic workflows, regression testing, benchmarking, and automated quality scoring.
  • Strong expertise with hybrid evaluation methodologies combining deterministic validation (citation grounding, structural validation, exact matching) and probabilistic LLM-as-a-Judge techniques (faithfulness, answer relevance, context precision, task completion).
  • Practical experience with RAG evaluation frameworks such as RAGAS, including evaluation tuning, embedding optimization, retrieval quality improvement, and production-scale LLM evaluation pipelines.
  • Experience validating hallucination, grounding, citation accuracy, bias, fairness, toxicity, and factual consistency in production LLM applications.
  • Experience designing systems that verify external knowledge claims and ensure responses are grounded in traceable citations and trusted data sources.
  • Strong experience testing cloud-native platforms and cloud-managed embedded products, including end-to-end system validation.
  • Experience establishing AI governance, safety, compliance, and Responsible AI practices for enterprise or safety-critical systems.
  • Proficiency in Python, Linux, shell scripting, modern test frameworks (PyTest, Playwright, Behave), and engineering productivity tools such as Jenkins and JIRA.
Ways to Stand Out:
  • Experience validating AI or agentic systems in safety-critical or regulated industries (automotive, aerospace, medical).
  • Track record building and scaling an AI test/evaluation platform or developer experience used by multiple teams (frameworks, reusable components, reference implementations).
  • Demonstrated wins moving AI testing practices from ad hoc to standardized, organization-wide adoption, with measurable impact on cycle time, quality, or reliability.
  • Experience implementing enterprise-grade AI governance (auditability, monitoring, policy enforcement) in production systems.
  • Deep experience evaluating LLM and RAG systems at scale, including agentic workflows, RAGAS-based evaluation, citation verification, hallucination detection, groundedness, faithfulness, answer relevance, tool/task correctness, and automated regression testing across offline and online feedback loops.

Sunnyvale HQ Benefits & Perks Offered:

  • Health care plan (Medical, Dental & Vision)
  • Flexible and Dependent Care Expense program
  • Retirement plan (401k)
  • Life Insurance (Basic, Voluntary & AD&D)
  • Unlimited paid time off per year, 14+ paid holidays
  • Hybrid office work arrangement
  • Complimentary lunches, snacks, and beverages during on-site working days
  • Wellness benefit allowance
  • Phone & Internet reimbursement
  • Computer Accessory Allowance