Senior AI Engineer
$170K - $200K/yr
Architect and maintain LLM-based systems, including prompt pipelines, agentic workflows, tool ... Contribute to internal standards for AI development: prompt management, evaluation frameworks ...
$170K - $200K/yr
Architect and maintain LLM-based systems, including prompt pipelines, agentic workflows, tool ... Contribute to internal standards for AI development: prompt management, evaluation frameworks ...
$170K - $200K/yr
Architect and maintain LLM-based systems, including prompt pipelines, agentic workflows, tool ... Contribute to internal standards for AI development: prompt management, evaluation frameworks ...
Seattle, WA · On-site
$170K - $200K/yr
Architect and maintain LLM-based systems, including prompt pipelines, agentic workflows, tool ... Contribute to internal standards for AI development: prompt management, evaluation frameworks ...
Seattle, WA · On-site
$170K - $200K/yr
Architect and maintain LLM-based systems, including prompt pipelines, agentic workflows, tool ... Contribute to internal standards for AI development: prompt management, evaluation frameworks ...
Seattle, WA · On-site
$170K - $200K/yr
Architect and maintain LLM-based systems, including prompt pipelines, agentic workflows, tool ... Contribute to internal standards for AI development: prompt management, evaluation frameworks ...
Seattle, WA · On-site
$170K - $200K/yr
Architect and maintain LLM-based systems, including prompt pipelines, agentic workflows, tool ... Contribute to internal standards for AI development: prompt management, evaluation frameworks ...
Conduct applied research to advance the state-of-the-art in LLM applications, including prompt engineering, few-shot learning, fine-tuning, and model evaluation • Production Deployment: Build ...
Conduct applied research to advance the state-of-the-art in LLM applications, including prompt engineering, few-shot learning, fine-tuning, and model evaluation • Production Deployment: Build ...
Seattle, WA · On-site
$190K - $219K/yr
... prompt management, evaluation frameworks, etc.) • Comfortable discussing architecture, APIs, infrastructure trade-offs, and scalability concerns with senior engineers • Strong analytical mindset ...
Seattle, WA · On-site
$190K - $219K/yr
... prompt management, evaluation frameworks, etc.) • Comfortable discussing architecture, APIs, infrastructure trade-offs, and scalability concerns with senior engineers • Strong analytical mindset ...
... prompt engineering and applied use cases such as grading, validation, or classification Strong working knowledge of evaluation methodology for generative AI, including LLM-as-a-judge design, meta ...
... prompt engineering and applied use cases such as grading, validation, or classification Strong working knowledge of evaluation methodology for generative AI, including LLM-as-a-judge design, meta ...
Position : Data Scientist with Generative AI (with AI , ML, and LLM - not ML Ops ) Location ... prompt tuning, retrieval augmentation (RAG), and model evaluation. * Optimize cost and performance ...
Quick apply
Position : Data Scientist with Generative AI (with AI , ML, and LLM - not ML Ops ) Location ... prompt tuning, retrieval augmentation (RAG), and model evaluation. * Optimize cost and performance ...
... prompt injection techniques • Explore edge cases to provoke disallowed, harmful, or incorrect ... or evaluation tooling • Comfort with structured data annotation and rubric-based scoring • ...
... prompt injection techniques • Explore edge cases to provoke disallowed, harmful, or incorrect ... or evaluation tooling • Comfort with structured data annotation and rubric-based scoring • ...
Bellevue, WA · On-site
$138K - $182K/yr
... with LLM services. Designing, testing, and optimizing prompts and prompt flows for LLMs and RAG ... evaluation metrics. Engineering Practices: TDD/unit integration tests, API documentation ...
Quick apply
Bellevue, WA · On-site
$138K - $182K/yr
... with LLM services. Designing, testing, and optimizing prompts and prompt flows for LLMs and RAG ... evaluation metrics. Engineering Practices: TDD/unit integration tests, API documentation ...
... production, builds the LLM integration layer, implements RAG pipelines, creates evaluation ... and optimize prompt engineering pipelines with version control, A/B testing, and regression ...
New
... production, builds the LLM integration layer, implements RAG pipelines, creates evaluation ... and optimize prompt engineering pipelines with version control, A/B testing, and regression ...
New
Seattle, WA · On-site
$130K - $156K/yr
... production, builds the LLM integration layer, implements RAG pipelines, creates evaluation ... and optimize prompt engineering pipelines with version control, A/B testing, and regression ...
New
Seattle, WA · On-site
$130K - $156K/yr
... production, builds the LLM integration layer, implements RAG pipelines, creates evaluation ... and optimize prompt engineering pipelines with version control, A/B testing, and regression ...
New
Bellevue, WA · On-site
$115K - $151K/yr
... prompt engineering, structured outputs (JSON schemas/function calling), and tool-augmented LLMs ... LLM evaluation using automated and human-in-the-loop techniques (offline + online). • Optimize ...
Bellevue, WA · On-site
$115K - $151K/yr
... prompt engineering, structured outputs (JSON schemas/function calling), and tool-augmented LLMs ... LLM evaluation using automated and human-in-the-loop techniques (offline + online). • Optimize ...
Seattle, WA · On-site
$32 - $95/hr
We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We've grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals ...
Quick apply
Seattle, WA · On-site
$32 - $95/hr
We work directly with frontier AI lab researchers to create evaluations, publish benchmarks, and push the boundary of data. We've grown from $0 to ~$1B run rate and pay ~$60M to over 30K individuals ...
Bellevue, WA · On-site
$64K - $98K/yr
LLM Orchestration; Prompt Engineering & AI Agents; Azure Cloud-Native Development & Microservices; Python; C#; Embeddings & AI Evaluation Frameworks; AI/ML Systems Development & Production Deployment.
Bellevue, WA · On-site
$64K - $98K/yr
LLM Orchestration; Prompt Engineering & AI Agents; Azure Cloud-Native Development & Microservices; Python; C#; Embeddings & AI Evaluation Frameworks; AI/ML Systems Development & Production Deployment.
This is a hands-on role across agentic workflows, RAG, prompt/policy design, LLM evaluation, and platform integration. You'll own the end-to-end path from use case evaluation production deployment ...
This is a hands-on role across agentic workflows, RAG, prompt/policy design, LLM evaluation, and platform integration. You'll own the end-to-end path from use case evaluation production deployment ...
Seattle, WA · Hybrid
... prompt/context patterns. * Implement LLM application patterns including RAG, document ingestion/chunking, embeddings, vector/hybrid search, and retrieval/evaluation telemetry. * Deliver governed ...
Seattle, WA · Hybrid
... prompt/context patterns. * Implement LLM application patterns including RAG, document ingestion/chunking, embeddings, vector/hybrid search, and retrieval/evaluation telemetry. * Deliver governed ...
Seattle, WA · On-site
$139K - $183K/yr
... LLM-powered features end-to-end. • Design and maintain evaluation frameworks for prompts ... prompt iteration. • Ability to move quickly and independently in a fast-paced environment. • ...
Seattle, WA · On-site
$139K - $183K/yr
... LLM-powered features end-to-end. • Design and maintain evaluation frameworks for prompts ... prompt iteration. • Ability to move quickly and independently in a fast-paced environment. • ...
Seattle, WA · On-site
... LLM-based agents for reasoning, synthesis, or evaluation tasks. • Familiarity with prompt engineering and multi-step reasoning--designing structured flows that balance quality, latency, and cost ...
Seattle, WA · On-site
... LLM-based agents for reasoning, synthesis, or evaluation tasks. • Familiarity with prompt engineering and multi-step reasoning--designing structured flows that balance quality, latency, and cost ...
Seattle, WA · On-site
$139K - $183K/yr
... maintain LLM-based systems, including prompt pipelines, agentic workflows, tool-calling ... management, evaluation frameworks, model selection, and system architecture Qualifications
Seattle, WA · On-site
$139K - $183K/yr
... maintain LLM-based systems, including prompt pipelines, agentic workflows, tool-calling ... management, evaluation frameworks, model selection, and system architecture Qualifications
Bellevue, WA · On-site
$155K - $167K/yr
Conduct LLM evaluation using automated and human-in-the-loop techniques (offline + online ... Prompt engineering expertise * Embeddings and vector search * Experienced in backend API design ...
Bellevue, WA · On-site
$155K - $167K/yr
Conduct LLM evaluation using automated and human-in-the-loop techniques (offline + online ... Prompt engineering expertise * Embeddings and vector search * Experienced in backend API design ...
$8.33 - $11.77
18% of jobs
$11.77 - $15.22
2% of jobs
$15.74 is the 25th percentile. Wages below this are outliers.
$15.22 - $18.66
30% of jobs
$18.66 - $22.11
15% of jobs
$22.11 - $25.55
9% of jobs
$26.07 is the 75th percentile. Wages above this are outliers.
$25.55 - $29
5% of jobs
$29 - $32.44
3% of jobs
$32.44 - $35.89
5% of jobs
$35.89 - $39.33
4% of jobs
$39.33 - $42.78
4% of jobs
$42.78 - $46.22
3% of jobs
$8
$23
$46
An LLM Prompt Evaluation job involves assessing and optimizing prompts used in large language models (LLMs) to ensure they generate accurate, relevant, and high-quality responses. Evaluators test different prompts, analyze model outputs, and refine phrasing to improve performance. This role requires a strong understanding of AI behavior, critical thinking, and sometimes domain-specific expertise to create effective instructions for the model.
Professionals in LLM Prompt Evaluation spend their days reviewing, analyzing, and scoring the outputs of large language models based on specific prompts. This involves identifying inaccuracies, biases, or other issues in AI-generated responses, and providing detailed, constructive feedback that informs further model development. Collaboration with data scientists, machine learning engineers, and product teams is common to align evaluation efforts with organizational goals. The role also includes documenting findings, participating in regular team meetings, and staying updated on best practices in prompt design and AI ethics.
To thrive in LLM Prompt Evaluation, you need a solid understanding of natural language processing, critical thinking, and analytical skills, often supported by a background in linguistics, computer science, or related disciplines. Familiarity with annotation tools, large language model (LLM) platforms, and prompt engineering frameworks is important. Attention to detail, strong written communication, and the ability to provide objective, structured feedback are key soft skills. These qualifications ensure accurate assessments of AI-generated responses and support the continuous improvement of language models.
The top searched job categories for Llm Prompt Evaluation jobs in Bothell, WA are:
Cities near Bothell, WA with the most Llm Prompt Evaluation job openings:

Seattle, WA
$170K - $200K/yr
Full-time
Medical, Life, Retirement
Re-posted 7 days ago
The real world is the next frontier, and at Metropolis, we are creating the artificial intelligence to make it responsive. We are pioneering the Recognition Economy - a future where mundane repetition disappears and being known unlocks access, comfort, and belonging everywhere you go. From transforming parking into a seamless drive-in, drive-out experience for millions of Members to expanding our intelligence layer across retail and hospitality, we are building a world that feels instinctive and magical. The future isn't coming; it's here, and we need builders, innovators, and problem solvers to help us create it.
Who you areMetropolis is seeking a Senior AI Engineer to join our Applied AI organization, a team purpose-built to rebuild how the company operates from an AI-first perspective. You will be the builder at the heart of this transformation, designing and shipping AI-powered tools and automation pipelines that replace manual, repetitive work across finance, operations, revenue, and beyond. This is a hands-on engineering role with enormous scope: you will own systems end-to-end, from the LLM prompt layer through to production deployment, working side-by-side with Process Analysts who understand the business deeply. If you want to see your work directly change how a company functions - not someday, but now - this is the role.
What you'll do4 Days in Office: Metropolis values in-person collaboration to drive innovation, strengthen culture, and enhance the Member experience. Our corporate team members hold to our office-first model, which requires employees to be on-site at least four days a week, fostering organic interactions that spark creativity and connection
When you join Metropolis, you'll join a team of world-class product leaders and engineers, building an ecosystem of technologies at the intersection of parking, mobility, and real estate. Our goal is to build an inclusive culture where everyone has a voice and the best idea wins. You will play a key role in building and maintaining this culture as our organization grows. The anticipated base salary for this position is $170,000.00 USD to $200,000.00 USD annually. The actual base salary offered is determined by a number of variables, including, as appropriate, the applicant's qualifications for the position, years of relevant experience, distinctive skills, level of education attained, certifications or other professional licenses held, and the location of residence and/or place of employment. Base salary is one component of Metropolis' total compensation package, which may also include access to or eligibility for healthcare benefits, a 401(k) plan, short-term and long-term disability coverage, basic life insurance, a lucrative stock option plan, bonus plans, and more. #LI-CM1 #LI-Onsite