The ideal candidate possesses a solid understanding of machine learning, natural language understanding, modern LLM architectures, LLM evaluation & tooling, and a passion for pushing boundaries in ...
The ideal candidate possesses a solid understanding of machine learning, natural language understanding, modern LLM architectures, LLM evaluation & tooling, and a passion for pushing boundaries in ...
The ideal candidate possesses a solid understanding of machine learning, natural language understanding, modern LLM architectures, LLM evaluation & tooling, and a passion for pushing boundaries in ...
The ideal candidate possesses a solid understanding of machine learning, natural language understanding, modern LLM architectures, LLM evaluation & tooling, and a passion for pushing boundaries in ...
Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time - RAG systems: Design retrieval-augmented generation systems that ...
Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time - RAG systems: Design retrieval-augmented generation systems that ...
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team
Seattle, WA · On-site
Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time - RAG systems: Design retrieval-augmented generation systems that ...
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team
Seattle, WA · On-site
Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time - RAG systems: Design retrieval-augmented generation systems that ...
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team
Seattle, WA · On-site
Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time - RAG systems: Design retrieval-augmented generation systems that ...
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team
Seattle, WA · On-site
Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time - RAG systems: Design retrieval-augmented generation systems that ...
Software Engineer III - Applied AI
Seattle, WA · On-site
$65.50 - $88/hr
Build and maintain LLM evaluation frameworks - including automated regression suites, human preference datasets, and task-specific benchmarks - to ensure consistent, trustworthy AI output * Architect ...
Software Engineer III - Applied AI
Seattle, WA · On-site
$65.50 - $88/hr
Build and maintain LLM evaluation frameworks - including automated regression suites, human preference datasets, and task-specific benchmarks - to ensure consistent, trustworthy AI output * Architect ...
Software Engineer III - Applied AI
Seattle, WA · On-site
$65.50 - $88/hr
Build and maintain LLM evaluation frameworks - including automated regression suites, human preference datasets, and task-specific benchmarks - to ensure consistent, trustworthy AI output * Architect ...
Software Engineer III - Applied AI
Seattle, WA · On-site
$65.50 - $88/hr
Build and maintain LLM evaluation frameworks - including automated regression suites, human preference datasets, and task-specific benchmarks - to ensure consistent, trustworthy AI output * Architect ...
Collaborate with partner teams on LLM evaluation frameworks and post-training methodologies. * Contribute to end-to-end delivery of solutions from research to production, including reusable science ...
Collaborate with partner teams on LLM evaluation frameworks and post-training methodologies. * Contribute to end-to-end delivery of solutions from research to production, including reusable science ...
MT/LLM Evaluation: Review machine-translated and AI-generated Turkish content for accuracy, fluency, and usability, determining whether content is publication-ready or requires post-editing.
Quick apply
MT/LLM Evaluation: Review machine-translated and AI-generated Turkish content for accuracy, fluency, and usability, determining whether content is publication-ready or requires post-editing.
Software Engineer III - Applied AI
Seattle, WA · On-site
$164.65 - $230.51/hr
Build and maintain LLM evaluation frameworks - including automated regression suites, human preference datasets, and task-specific benchmarks - to ensure consistent, trustworthy AI output* Architect ...
Software Engineer III - Applied AI
Seattle, WA · On-site
$164.65 - $230.51/hr
Build and maintain LLM evaluation frameworks - including automated regression suites, human preference datasets, and task-specific benchmarks - to ensure consistent, trustworthy AI output* Architect ...
Lead AI Engineer
Bellevue, WA · On-site
$115K - $151K/yr
... • Conduct LLM evaluation using automated and human-in-the-loop techniques (offline + online). • Optimize inference workflows for latency, GPU utilization, and cost efficiency (quantization ...
Lead AI Engineer
Bellevue, WA · On-site
$115K - $151K/yr
... • Conduct LLM evaluation using automated and human-in-the-loop techniques (offline + online). • Optimize inference workflows for latency, GPU utilization, and cost efficiency (quantization ...
Principal AI Engineer
Seattle, WA · On-site
Define's Client enterprise AI strategy and architecture, including platform selection, LLM evaluation, governance model, technical standards, and alignment with Client's long-term digital and ...
Quick apply
Principal AI Engineer
Seattle, WA · On-site
Define's Client enterprise AI strategy and architecture, including platform selection, LLM evaluation, governance model, technical standards, and alignment with Client's long-term digital and ...
Principal AI Engineer
Seattle, WA · On-site
Define's Client enterprise AI strategy and architecture, including platform selection, LLM evaluation, governance model, technical standards, and alignment with Client's long-term digital and ...
Principal AI Engineer
Seattle, WA · On-site
Define's Client enterprise AI strategy and architecture, including platform selection, LLM evaluation, governance model, technical standards, and alignment with Client's long-term digital and ...
Apple Services Engineering (ASE) powers the AI and LLM features behind experiences that hundreds of ... As these systems increasingly rely on human-in-the-loop evaluation, the quality of our products is ...
Apple Services Engineering (ASE) powers the AI and LLM features behind experiences that hundreds of ... As these systems increasingly rely on human-in-the-loop evaluation, the quality of our products is ...
Deployed Engineer (Seattle)
Seattle, WA · On-site
... Worked with LLM evaluation, observability, or guardrails • Have experience with cloud environments (AWS, GCP, Azure), containers, and basic Kubernetes concepts • Have shipped and operated ...
Deployed Engineer (Seattle)
Seattle, WA · On-site
... Worked with LLM evaluation, observability, or guardrails • Have experience with cloud environments (AWS, GCP, Azure), containers, and basic Kubernetes concepts • Have shipped and operated ...
Apple Services Engineering (ASE) powers the AI and LLM features behind experiences that hundreds of ... As these systems increasingly rely on human-in-the-loop evaluation, the quality of our products is ...
Apple Services Engineering (ASE) powers the AI and LLM features behind experiences that hundreds of ... As these systems increasingly rely on human-in-the-loop evaluation, the quality of our products is ...
Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities. * Apply post-training expertise (SFT, RLHF, reward modeling) to connect ...
Design and build benchmarks and evaluation methods that measure LLM capabilities in both text and multimodal modalities. * Apply post-training expertise (SFT, RLHF, reward modeling) to connect ...
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team (Seattle)
Seattle, WA · On-site
LLM evaluation: Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time * RAG systems: Design retrieval-augmented ...
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team (Seattle)
Seattle, WA · On-site
LLM evaluation: Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time * RAG systems: Design retrieval-augmented ...
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team (Seattle)
Seattle, WA · On-site
LLM evaluation: Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time * RAG systems: Design retrieval-augmented ...
New
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team (Seattle)
Seattle, WA · On-site
LLM evaluation: Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time * RAG systems: Design retrieval-augmented ...
New
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team
Seattle, WA · On-site
LLM evaluation:** Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time - **RAG systems:** Design retrieval-augmented ...
Senior Applied Scientist, Amazon Global Data Center Ops Central Insight and Analytics Team
Seattle, WA · On-site
LLM evaluation:** Build evaluation frameworks that measure LLM output quality (accuracy, actionability, calibration) and detect degradation over time - **RAG systems:** Design retrieval-augmented ...
Llm Evaluation information
What is the difference between Llm Evaluation vs Data Scientist?
| Aspect | Llm Evaluation | Data Scientist |
|---|---|---|
| Required Credentials | Typically requires knowledge of machine learning, NLP, and AI concepts; often a degree in computer science or related fields | Requires degrees in computer science, statistics, or related fields; often includes certifications in data analysis or machine learning |
| Work Environment | Primarily research and testing environments, focusing on AI model assessment | Data analysis, modeling, and visualization in various industries like finance, healthcare, or tech |
| Employer & Industry Usage | Used by AI research labs, tech companies, and organizations developing NLP models | Used across industries for data analysis, predictive modeling, and business insights |
While both roles involve working with data and machine learning, Llm Evaluation focuses on assessing large language models' performance, whereas Data Scientists develop and implement data-driven solutions across various sectors.
What are popular job titles related to Llm Evaluation jobs in Bothell, WA?
For Llm Evaluation jobs in Bothell, WA, the most frequently searched job titles are:
What job categories do people searching Llm Evaluation jobs in Bothell, WA look for?
The top searched job categories for Llm Evaluation jobs in Bothell, WA are:
What cities near Bothell, WA are hiring for Llm Evaluation jobs?
Cities near Bothell, WA with the most Llm Evaluation job openings:

Full-time
Re-posted 21 days ago
Amazon rating
7.4
Based on 7,147 frontline employees who took The Breakroom Quiz
5th of 39 rated national retailers
Job description
Alexa International is looking for a passionate, talented, and inventive Applied Scientist to help build industry-leading technology with Large Language Models (LLMs) and multimodal systems, requiring strong deep learning and generative models knowledge. You will contribute to developing novel solutions and deliver high-quality results that impact Alexa's international products and services.
Key job responsibilities
As an Applied Scientist with the Alexa International team, you will work with talented peers to develop novel algorithms and modeling techniques to advance the state of the art with LLMs. Your work will directly impact our international customers in the form of products and services that make use of digital assistant technology
You will leverage Amazon's heterogeneous data sources, unique and diverse international customer nuances and large-scale computing resources to accelerate advances in text, voice, and vision domains in a multimodal setup. The ideal candidate possesses a solid understanding of machine learning, natural language understanding, modern LLM architectures, LLM evaluation & tooling, and a passion for pushing boundaries in this vast and quickly evolving field. They thrive in fast-paced environments to tackle complex challenges, excel at swiftly delivering impactful solutions while iterating based on user feedback, and collaborate effectively with cross-functional teams.
A day in the life
* Analyze, understand, and model customer behavior and the customer experience based on large-scale data.
* Build novel online & offline evaluation metrics and methodologies for multimodal personal digital assistants.
* Fine-tune/post-train LLMs using techniques like SFT, DPO, RLHF, and RLAIF.
* Set up experimentation frameworks for agile model analysis and A/B testing.
* Collaborate with partner teams on LLM evaluation frameworks and post-training methodologies.
* Contribute to end-to-end delivery of solutions from research to production, including reusable science components.
* Communicate solutions clearly to partners and stakeholders.
* Contribute to the scientific community through publications and community engagement.
About Amazon
Sourced by ZipRecruiter
Amazon.com, Inc., commonly known as Amazon, is an American multinational technology company. It was founded by Jeff Bezos in 1994 and initially started as an online marketplace for books. Since then, Amazon has expanded its operations and become one of the largest e-commerce companies in the world. Amazon's primary business is its online retail platform, where customers can purchase a vast array of products, including electronics, clothing, books, home goods, and much more. The company offers a convenient and user-friendly shopping experience, with features such as fast shipping, customer reviews, and personalized recommendations. In addition to its e-commerce platform, Amazon has diversified its business into various other areas. One of its notable ventures is Amazon Web Services (AWS), a comprehensive cloud computing platform that provides services such as storage, compute power, and database management to individuals and businesses. AWS has become a leader in the cloud computing industry, powering many websites and applications worldwide. Amazon has also developed its own consumer electronics, including the popular Amazon Kindle e-reader, Fire tablets, Fire TV streaming devices, and the Alexa-powered Echo smart speakers. The Alexa voice assistant, integrated into these devices, allows users to interact with their devices using voice commands, perform tasks, and access information. Furthermore, Amazon has expanded into media and entertainment. It operates Prime Video, a streaming service that offers a wide range of movies, TV shows, and original content. Amazon Music provides a platform for streaming and purchasing digital music, while Audible offers audiobooks and other audio content. The company's commitment to customer satisfaction and convenience is demonstrated by its membership program, Amazon Prime. Prime members receive various benefits, including free two-day shipping, access to streaming services, exclusive deals, and more.
Industry
It services, book publishers, retail, real estate, computer and electronic product manufacturing and software development
Company size
10,000+ Employees
Headquarters location
Seattle, WA, US