... LLM evaluation science for AI-powered agent experiences, or agent memory architectures that enable persistent, adaptive behavior across sessions. You will work closely with applied scientists ...
... LLM evaluation science for AI-powered agent experiences, or agent memory architectures that enable persistent, adaptive behavior across sessions. You will work closely with applied scientists ...
... LLM evaluation science for AI-powered agent experiences, or agent memory architectures that enable persistent, adaptive behavior across sessions. You will work closely with applied scientists ...
... LLM evaluation science for AI-powered agent experiences, or agent memory architectures that enable persistent, adaptive behavior across sessions. You will work closely with applied scientists ...
Lead AI Engineer
Bellevue, WA · On-site
$155K - $167K/yr
Conduct LLM evaluation using automated and human-in-the-loop techniques (offline + online). * Optimize inference workflows for latency , GPU utilization , and cost efficiency (quantization, batching ...
Lead AI Engineer
Bellevue, WA · On-site
$155K - $167K/yr
Conduct LLM evaluation using automated and human-in-the-loop techniques (offline + online). * Optimize inference workflows for latency , GPU utilization , and cost efficiency (quantization, batching ...
Technical Product Manager, LLM/ML Domain
Seattle, WA · On-site
$190K - $219K/yr
... evaluation, retraining, versioning, and observability of ML/LLM systems • Drive adoption through documentation, training, and internal evangelism • Collaborate with engineering to level up Agoda ...
Technical Product Manager, LLM/ML Domain
Seattle, WA · On-site
$190K - $219K/yr
... evaluation, retraining, versioning, and observability of ML/LLM systems • Drive adoption through documentation, training, and internal evangelism • Collaborate with engineering to level up Agoda ...
Deployed Engineer (Seattle)
Seattle, WA · On-site
$165K - $380K/yr
Worked with LLM evaluation, observability, or guardrails * Have experience with cloud environments (AWS, GCP, Azure), containers, and basic Kubernetes concepts * Have shipped and operated production ...
Deployed Engineer (Seattle)
Seattle, WA · On-site
$165K - $380K/yr
Worked with LLM evaluation, observability, or guardrails * Have experience with cloud environments (AWS, GCP, Azure), containers, and basic Kubernetes concepts * Have shipped and operated production ...
Design rigorous evaluation frameworks to measure model performance, bias, safety, and business ... LLM applications (prompt engineering, RAG, fine-tuning, RLHF, Transformer models, and AI Agents ...
Design rigorous evaluation frameworks to measure model performance, bias, safety, and business ... LLM applications (prompt engineering, RAG, fine-tuning, RLHF, Transformer models, and AI Agents ...
Principle AI Solution Architect - Healthcare
Seattle, WA · On-site
$71.75 - $94.50/hr
Define's Client enterprise AI strategy and architecture, including platform selection, LLM evaluation, governance model, technical standards, and alignment with Client's long-term digital and ...
Principle AI Solution Architect - Healthcare
Seattle, WA · On-site
$71.75 - $94.50/hr
Define's Client enterprise AI strategy and architecture, including platform selection, LLM evaluation, governance model, technical standards, and alignment with Client's long-term digital and ...
Principle AI Solution Architect - Healthcare
Seattle, WA · On-site
$71.75 - $94.50/hr
Define's Client enterprise AI strategy and architecture, including platform selection, LLM evaluation, governance model, technical standards, and alignment with Client's long-term digital and ...
Quick apply
Principle AI Solution Architect - Healthcare
Seattle, WA · On-site
$71.75 - $94.50/hr
Define's Client enterprise AI strategy and architecture, including platform selection, LLM evaluation, governance model, technical standards, and alignment with Client's long-term digital and ...
QA Reviewer & Project Coordinator (On-Site - Seattle or Boston)
Seattle, WA · On-site
$90 - $140/hr
Experience with AI safety, LLM evaluation, prompt engineering, or red teaming. * Background in trust & safety, content moderation, or behavioral policy development. * Experience developing evaluation ...
QA Reviewer & Project Coordinator (On-Site - Seattle or Boston)
Seattle, WA · On-site
$90 - $140/hr
Experience with AI safety, LLM evaluation, prompt engineering, or red teaming. * Background in trust & safety, content moderation, or behavioral policy development. * Experience developing evaluation ...
AI Automation Engineer L1
Bellevue, WA · On-site
$100 - $180/hr
Build LLM evaluation pipelines covering relevance, groundedness, and hallucination metrics integrated into CI/CD workflows to enforce automated quality gates before every model or agent deployment.
AI Automation Engineer L1
Bellevue, WA · On-site
$100 - $180/hr
Build LLM evaluation pipelines covering relevance, groundedness, and hallucination metrics integrated into CI/CD workflows to enforce automated quality gates before every model or agent deployment.
AI Engineer
$50K - $112K/yr
... LLM evaluation, and fine-tuning, to develop production-ready applications powered by foundation models - Developing automated evaluation frameworks, including LLM-as-judge pipelines, regression ...
AI Engineer
$50K - $112K/yr
... LLM evaluation, and fine-tuning, to develop production-ready applications powered by foundation models - Developing automated evaluation frameworks, including LLM-as-judge pipelines, regression ...
Experience with AI safety, LLM evaluation, prompt engineering, or red teaming. * Background in trust & safety, content moderation, or behavioral policy development. * Experience developing evaluation ...
Experience with AI safety, LLM evaluation, prompt engineering, or red teaming. * Background in trust & safety, content moderation, or behavioral policy development. * Experience developing evaluation ...
Experience with AI safety, LLM evaluation, prompt engineering, or red teaming. * Background in trust & safety, content moderation, or behavioral policy development. * Experience developing evaluation ...
Quick apply
Experience with AI safety, LLM evaluation, prompt engineering, or red teaming. * Background in trust & safety, content moderation, or behavioral policy development. * Experience developing evaluation ...
AIML - Sr Manager, Evaluation - Data Science & Insights
Seattle, WA · On-site
$225K - $381K/yr
Driving hands-on execution across human evaluation, and LLM-based autograders and rubrics, and ensuring those methods translate into measurable model and agentic feature improvements. Partnering with ...
AIML - Sr Manager, Evaluation - Data Science & Insights
Seattle, WA · On-site
$225K - $381K/yr
Driving hands-on execution across human evaluation, and LLM-based autograders and rubrics, and ensuring those methods translate into measurable model and agentic feature improvements. Partnering with ...
Research Scientist, AI Evaluation Science
$205K - $308K/yr
LLM-as-judge calibration, rubric design, and bias detection; intelligent evaluation strategies including active learning for test selection and automated failure discovery; or validity frameworks for ...
Research Scientist, AI Evaluation Science
$205K - $308K/yr
LLM-as-judge calibration, rubric design, and bias detection; intelligent evaluation strategies including active learning for test selection and automated failure discovery; or validity frameworks for ...
AIML - Sr Manager, Evaluation - Data Science & Insights
Seattle, WA · On-site
$180 - $240/hr
Successful candidates will have strong experience in traditional human evaluation methodology, in addition to hands-on experience building and deploying LLM-based autograders and rubrics, and using ...
New
AIML - Sr Manager, Evaluation - Data Science & Insights
Seattle, WA · On-site
$180 - $240/hr
Successful candidates will have strong experience in traditional human evaluation methodology, in addition to hands-on experience building and deploying LLM-based autograders and rubrics, and using ...
New
QA Reviewer & Project Coordinator (On-Site - Seattle or Boston)
Seattle, WA · On-site
$110K - $125K/yr
Experience with AI safety, LLM evaluation, prompt engineering, or red teaming. * Background in trust & safety, content moderation, or behavioral policy development. * Experience developing evaluation ...
QA Reviewer & Project Coordinator (On-Site - Seattle or Boston)
Seattle, WA · On-site
$110K - $125K/yr
Experience with AI safety, LLM evaluation, prompt engineering, or red teaming. * Background in trust & safety, content moderation, or behavioral policy development. * Experience developing evaluation ...
AIML - Sr Manager, Evaluation - Data Science & Insights
Seattle, WA · On-site
$226 - $382/hr
Driving hands-on execution across human evaluation, and LLM-based autograders and rubrics, and ensuring those methods translate into measurable model and agentic feature improvements. * Partnering ...
AIML - Sr Manager, Evaluation - Data Science & Insights
Seattle, WA · On-site
$226 - $382/hr
Driving hands-on execution across human evaluation, and LLM-based autograders and rubrics, and ensuring those methods translate into measurable model and agentic feature improvements. * Partnering ...
... LLM evaluation science for AI-powered agent experiences, or agent memory architectures that enable persistent, adaptive behavior across sessions. You will work closely with applied scientists ...
... LLM evaluation science for AI-powered agent experiences, or agent memory architectures that enable persistent, adaptive behavior across sessions. You will work closely with applied scientists ...
Senior Program Manager, Human Data Operations
Seattle, WA · On-site
$132K - $132K/yr
Experience with AI safety, LLM evaluation, human-in-the-loop systems, annotation, red teaming, or trust & safety. * Experience leading multi-site operations or facilities coordination. * Experience ...
Senior Program Manager, Human Data Operations
Seattle, WA · On-site
$132K - $132K/yr
Experience with AI safety, LLM evaluation, human-in-the-loop systems, annotation, red teaming, or trust & safety. * Experience leading multi-site operations or facilities coordination. * Experience ...
Llm Evaluation information
What is the difference between Llm Evaluation vs Data Scientist?
| Aspect | Llm Evaluation | Data Scientist |
|---|---|---|
| Required Credentials | Typically requires knowledge of machine learning, NLP, and AI concepts; often a degree in computer science or related fields | Requires degrees in computer science, statistics, or related fields; often includes certifications in data analysis or machine learning |
| Work Environment | Primarily research and testing environments, focusing on AI model assessment | Data analysis, modeling, and visualization in various industries like finance, healthcare, or tech |
| Employer & Industry Usage | Used by AI research labs, tech companies, and organizations developing NLP models | Used across industries for data analysis, predictive modeling, and business insights |
While both roles involve working with data and machine learning, Llm Evaluation focuses on assessing large language models' performance, whereas Data Scientists develop and implement data-driven solutions across various sectors.
What are popular job titles related to Llm Evaluation jobs in Bothell, WA?
For Llm Evaluation jobs in Bothell, WA, the most frequently searched job titles are:
What job categories do people searching Llm Evaluation jobs in Bothell, WA look for?
The top searched job categories for Llm Evaluation jobs in Bothell, WA are:
What cities near Bothell, WA are hiring for Llm Evaluation jobs?
Cities near Bothell, WA with the most Llm Evaluation job openings:

Amazon rating
7.4
Based on 7,147 frontline employees who took The Breakroom Quiz
5th of 39 rated national retailers
Job description
AWS Insights & Optimization is looking for an Applied Scientist to help develop sophisticated algorithms and models that involve analyzing and learning from over 540 billion customer cost, usage, and utilization events daily. We use this data to power cost anomaly detection, cost forecasting, rightsizing recommendations across seven compute services, Savings Plans optimization, and AI-powered conversational experiences that help customers understand and optimize their AWS spend. Our team's vision is to be the world's provider of intelligent AWS cloud financial management, where customers can understand, control, and optimize usage of AWS products.
We sit at the nexus of all AWS services and interact directly with end-customers, building relationships with teams across AWS to ensure we offer a secure and reliable experience that builds trust and provides intelligent insights.
As a successful Applied Scientist in AWS Insights & Optimization, you will own models end-to-end - from problem formulation through experimentation to production deployment. Your work may span cost anomaly detection (decomposition, detection, and root cause attribution), time-series forecasting with a focus on accuracy and consistency, rightsizing engines for EC2, EBS, Lambda, ECS, RDS, and Aurora, LLM evaluation science for AI-powered agent experiences, or agent memory architectures that enable persistent, adaptive behavior across sessions. You will work closely with applied scientists, software engineers, and product teams to enhance existing models and build new ones that solve challenging customer problems
You will drive implementation of proposed models, establish testing strategies to validate them before and after production, and define evaluation metrics that determine whether capabilities meet the quality bar. We value accuracy over speed, measurability over intuition, and simplicity over complexity - the simplest model that meets the bar wins. You are an analytical problem solver who enjoys diving into data, are excited about investigating and developing algorithms, and can influence technical teams and business stakeholders to solve real-world customer problems.
About Amazon
Sourced by ZipRecruiter
Amazon.com, Inc., commonly known as Amazon, is an American multinational technology company. It was founded by Jeff Bezos in 1994 and initially started as an online marketplace for books. Since then, Amazon has expanded its operations and become one of the largest e-commerce companies in the world. Amazon's primary business is its online retail platform, where customers can purchase a vast array of products, including electronics, clothing, books, home goods, and much more. The company offers a convenient and user-friendly shopping experience, with features such as fast shipping, customer reviews, and personalized recommendations. In addition to its e-commerce platform, Amazon has diversified its business into various other areas. One of its notable ventures is Amazon Web Services (AWS), a comprehensive cloud computing platform that provides services such as storage, compute power, and database management to individuals and businesses. AWS has become a leader in the cloud computing industry, powering many websites and applications worldwide. Amazon has also developed its own consumer electronics, including the popular Amazon Kindle e-reader, Fire tablets, Fire TV streaming devices, and the Alexa-powered Echo smart speakers. The Alexa voice assistant, integrated into these devices, allows users to interact with their devices using voice commands, perform tasks, and access information. Furthermore, Amazon has expanded into media and entertainment. It operates Prime Video, a streaming service that offers a wide range of movies, TV shows, and original content. Amazon Music provides a platform for streaming and purchasing digital music, while Audible offers audiobooks and other audio content. The company's commitment to customer satisfaction and convenience is demonstrated by its membership program, Amazon Prime. Prime members receive various benefits, including free two-day shipping, access to streaming services, exclusive deals, and more.
Industry
It services, book publishers, retail, real estate, computer and electronic product manufacturing and software development
Company size
10,000+ Employees
Headquarters location
Seattle, WA, US