1

Llm Prompt Evaluation Jobs (NOW HIRING)

GenAI Architect / AI Architect (LLM, Prompt Engineering, RAG, AWS/Azure) Location: Torrance, CA ... Implement evaluation frameworks to measure LLM performance (accuracy, hallucination, latency ...

AI Evaluation Scientist

Mclean, VA · On-site

$105K - $145K/yr

Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures. * Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness ...

Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures. * Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness ...

AI Evaluation Scientist

Mclean, VA · On-site

$105K - $145K/yr

Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures. * Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness ...

AI Evaluation Scientist

Mclean, VA · On-site

$105K - $145K/yr

Understanding of LLM behavior, prompt evaluation, retrieval pipelines, or RAG architectures. * Exposure to responsible AI concepts and governance-aligned evaluation criteria (e.g., fairness ...

Showing results 21-40

Llm Prompt Evaluation information

See salary details

$7

$20

$41

How much do llm prompt evaluation jobs pay per hour?

As of Aug 8, 2026, the average hourly pay for llm prompt evaluation in the United States is $20.89, according to ZipRecruiter salary data. Most workers in this role earn between $14.42 and $27.64 per hour, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive in the LLM prompt evaluation position?

To thrive in LLM Prompt Evaluation, you need a solid understanding of natural language processing, critical thinking, and analytical skills, often supported by a background in linguistics, computer science, or related disciplines. Familiarity with annotation tools, large language model (LLM) platforms, and prompt engineering frameworks is important. Attention to detail, strong written communication, and the ability to provide objective, structured feedback are key soft skills. These qualifications ensure accurate assessments of AI-generated responses and support the continuous improvement of language models.

What is an LLM prompt evaluation?

An LLM Prompt Evaluation job involves assessing and optimizing prompts used in large language models (LLMs) to ensure they generate accurate, relevant, and high-quality responses. Evaluators test different prompts, analyze model outputs, and refine phrasing to improve performance. This role requires a strong understanding of AI behavior, critical thinking, and sometimes domain-specific expertise to create effective instructions for the model.

What does an LLM prompt evaluation do?

Professionals in LLM Prompt Evaluation spend their days reviewing, analyzing, and scoring the outputs of large language models based on specific prompts. This involves identifying inaccuracies, biases, or other issues in AI-generated responses, and providing detailed, constructive feedback that informs further model development. Collaboration with data scientists, machine learning engineers, and product teams is common to align evaluation efforts with organizational goals. The role also includes documenting findings, participating in regular team meetings, and staying updated on best practices in prompt design and AI ethics.

More about Llm Prompt Evaluation jobs
What cities are hiring for Llm Prompt Evaluation jobs? Cities with the most Llm Prompt Evaluation job openings:
What states have the most Llm Prompt Evaluation jobs? States with the most job openings for Llm Prompt Evaluation jobs include:
Infographic showing various Llm Prompt Evaluation job openings in the United States as of August 2026, with employment types broken down into 2% As Needed, 79% Full Time, 17% Part Time, and 2% Contract. Highlights an 90% Physical, 2% Hybrid, and 8% Remote job distribution, with an average salary of $43,449 per year, or $20.9 per hour.

Principal Engineer - Context Engineering & LLM Optimization

Bank of America

Charlotte, NC • On-site

$140 - $180/hr

Other

Re-posted 9 days ago


Bank Of America rating

8.2

Company rating: 8.2 out of 10

Based on 525 frontline employees who took The Breakroom Quiz

51st of 170 rated banks


Job description

The Context Engineering & LLM Optimization Principal Engineer is responsible for defining and leading the engineering approach for solutions at the program or portfolio level, to deliver significant business outcomes. Key responsibilities include continuously improving the design, quality, and reuse of the solution and delivering technology enablers that improve development efficiencies for the solution. Job expectations include familiarity with at least one area of engineering, acting as a “go to” reference across the organization, and applying knowledge to improve technical competencies through recruitment and development activities.

Developer Experience (DevEx) provides enterprise technical standards and common technical services, platforms, and tools that are leveraged by delivery teams across all lines of business. Within the SDLC Software Delivery Lifecycle program, this role leads portfolio product delivery strategy and execution for enterprise software delivery capabilities, ensuring the right investments, operating model, governance, and prioritization are in place to improve how internal technical users build, test, and deliver software at scale.

The Context Engineering & LLM Optimization Principal Engineer is responsible for designing how information is selected, organized, compressed, prioritized, and presented to LLMs. This role focuses on context window management, prompt architecture, retrieval orchestration, grounding strategies, instruction design, tool‑use patterns, and evaluation of LLM behavior.

This engineer ensures that LLM applications receive the right context, in the right format, at the right time, with minimal token waste and maximum answer quality.

Responsibilities
  • Develops the engineering approach for the entire program/portfolio solution and works with Architecture, to develop/analyze/deliver the implementation of technical enablers
  • Leads the planning, definition, and design of the complex features which span multiple teams and explore solution alternatives
  • Creates ideas on designing complex technology and solution development approaches
  • Leads the technical oversight for teams in solution development including design reviews and code within own domain
  • Defines the technology tool stack for the solution within ranged of internally approved and supported technologies
  • Explores state-of-the-art technologies to improve development efficiencies, quality of test/QA coverage, and release management
  • Leads and is responsible for the end-to-end test strategy/creation/adherence, and the integration between teams for a program/portfolio solution
  • Design context engineering strategies for enterprise LLM and RAG applications.
  • Define prompt architectures for system prompts, developer instructions, user prompts, retrieved context, tool outputs, conversation history, and structured constraints.
  • Optimize context window usage through summarization, compression, ranking, filtering, deduplication, and context prioritization.
  • Design retrieval orchestration patterns that determine what data is retrieved, when it is retrieved, and how it is injected into the LLM prompt.
  • Partner with RAG database engineers to tune retrieval outputs for downstream reasoning quality.
  • Partner with data ingestion engineers to improve source formatting, metadata, and chunk structures for better contextual use.
  • Develop patterns for multi-turn conversation memory, session state, user intent preservation, and context refresh.
  • Define strategies for grounding, citation handling, source attribution, conflicting evidence resolution, and hallucination reduction.
  • Improve the experience for our developers, making it easier to deliver industry-leading solutions, while managing work efficiently and with the right controls
  • Advance our technology platforms through innovation
  • Reduce risk and improve quality across our technology portfolio by aligning to a single enterprise architecture strategy and delivering governance that enables consistency, integration and automation
  • Design LLM evaluation frameworks for answer quality, factuality, instruction adherence, relevance, safety, and token efficiency.
  • Establish prompt engineering and context engineering standards across product and platform teams.
  • Evaluate LLM model behavior across different context sizes, retrieval strategies, and prompt structures.
  • Define reusable patterns for agents, tool calling, function calling, dynamic prompt generation, and workflow-based reasoning.
  • Lead technical reviews for LLM application design, prompt safety, and context efficiency.
  • Serve as a senior technical authority for enterprise AI platform engineering.
  • Own architecture decisions that impact multiple teams, systems, or domains.
  • Create reusable patterns, reference architectures, standards, and engineering guardrails.
  • Mentor senior engineers and influence technical direction without requiring direct reporting authority.
  • Balance innovation with operational reliability, security, compliance, scalability, and cost management.
  • Communicate complex AI and data engineering concepts clearly to engineering, product, risk, security, and executive stakeholders.
Required Qualifications
  • 10+ years of software engineering, data engineering, platform engineering, or AI engineering experience.
  • 5+ years designing large-scale enterprise systems.
  • 2+ years working with LLM, RAG, vector search, semantic search, or AI platform capabilities.
  • Experience operating systems in regulated, security-conscious, or enterprise-scale environments.
  • Extensive experience building or architecting production LLM, RAG, or AI assistant systems.
  • Deep understanding of how LLMs use prompts, retrieved context, conversation history, system instructions, and tool outputs.
  • Strong knowledge of context window management, token budgeting, prompt construction, grounding, and response evaluation.
  • Experience with OpenAI, Azure OpenAI, Anthropic, Google Gemini, Meta Llama, or similar LLM ecosystems.
  • Experience designing prompt templates, retrieval‑augmented prompts, agent workflows, and tool‑use orchestration.
  • Familiarity with vector search, embeddings, reranking, semantic retrieval, and document chunking.
  • Experience with automated LLM evaluation, prompt regression testing, and quality measurement.
  • Ability to define enterprise standards for reliable, explainable, and secure LLM behavior.
  • Proven ability to lead architecture across multiple engineering teams.
  • Strong written and verbal communication skills.
  • Bachelor’s degree in Computer Science, Engineering, Information Systems, Applied Mathematics, or a related technical field
Desired Qualifications
  • Experience with agentic workflows, multi-agent orchestration, function calling, or tool‑augmented reasoning.
  • Experience with prompt injection mitigation, jailbreak resistance, and secure context handling.
  • Experience with token optimization, long‑context models, summarization pipelines, and contextual compression.
  • Experience with user personalization, enterprise memory patterns, or domain‑specific copilots.
  • Higher‑quality LLM responses with better grounding and reduced hallucination.
  • Lower token usage and improved response latency through efficient context construction.
  • Standardized prompt and context patterns reused across teams.
  • Improved evaluation coverage for LLM behavior, factuality, and instruction adherence.
  • Better alignment between retrieved enterprise knowledge and generated responses.
  • Enterprise architecture
  • Distributed systems design
  • AI platform engineering
  • Data governance and security
  • Cloud‑native engineering
  • Observability and operational excellence
  • Technical strategy and roadmap development
  • Cross‑functional influence
  • Vendor and platform evaluation
  • Production support and continuous improvement
Skills
  • Automation
  • Influence
  • Result Orientation
  • Stakeholder Management
  • Technical Strategy Development
  • Application Development
  • Architecture
  • Business Acumen
  • Risk Management
  • Solution Design
  • Agile Practices
  • Analytical Thinking
  • Collaboration
  • Data Management
  • Solution Delivery Process

Shift: 1st shift (United States of America)

Hours Per Week: 40

#J-18808-Ljbffr

What Bank Of America employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Bank Of America logo

About Bank Of America

Sourced by ZipRecruiter

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. Responsible Growth is how we run our company and how we deliver for our clients, teammates, communities and shareholders every day. One of the keys to driving Responsible Growth is being a great place to work for our teammates around the world. We're devoted to being a diverse and inclusive workplace for everyone. We hire individuals with a broad range of backgrounds and experiences and invest heavily in our teammates and their families by offering competitive benefits to support their physical, emotional, and financial well-being.

Industry

Finance and insurance

Company size

10,000+ Employees

Headquarters location

Charlotte, NC, US

Year founded

1998

Social media