LLM Optimization & Evaluation
Austin, TX · On-site
Info Way Solutions is seeking a highly skilled LLM Optimization & Evaluation Specialist to design, optimize, and evaluate large language model systems in production. This role focuses on prompt ...
Austin, TX · On-site
Info Way Solutions is seeking a highly skilled LLM Optimization & Evaluation Specialist to design, optimize, and evaluate large language model systems in production. This role focuses on prompt ...
Austin, TX · On-site
Info Way Solutions is seeking a highly skilled LLM Optimization & Evaluation Specialist to design, optimize, and evaluate large language model systems in production. This role focuses on prompt ...
Columbia, MD · On-site +1
$93K - $100K/yr
Familiarity with LLM evaluation frameworks, structured benchmarking, or human-in-the-loop refinement methods (e.g., RLHF-style workflows). * Expertise with advanced retrieval techniques such as ...
Columbia, MD · On-site +1
$93K - $100K/yr
Familiarity with LLM evaluation frameworks, structured benchmarking, or human-in-the-loop refinement methods (e.g., RLHF-style workflows). * Expertise with advanced retrieval techniques such as ...
Santa Clara, CA · On-site
$73.50 - $96.75/hr
Mentor engineering teams on LLM architecture, evaluation, performance tuning, and Agentic AI development. Required Education Bachelor's or Master's degree in Computer Science, Data Science ...
Santa Clara, CA · On-site
$73.50 - $96.75/hr
Mentor engineering teams on LLM architecture, evaluation, performance tuning, and Agentic AI development. Required Education Bachelor's or Master's degree in Computer Science, Data Science ...
Evanston, IL · On-site
Experience with LLM evaluation frameworks and metric-driven iteration.
Evanston, IL · On-site
Experience with LLM evaluation frameworks and metric-driven iteration.
Austin, TX · On-site
LLM Evaluation Frameworks & Responsible AI Practices * REST API Development & Automated API Testing * Java/Kotlin Development * CI/CD using Maven & Jenkins * Kubernetes Container Orchestration
Quick apply
Austin, TX · On-site
LLM Evaluation Frameworks & Responsible AI Practices * REST API Development & Automated API Testing * Java/Kotlin Development * CI/CD using Maven & Jenkins * Kubernetes Container Orchestration
Master's degree or PhD in Engineering, Computer Science, or a related technical field. * 5 years of experience with LLM evaluation, machine learning algorithms and tools, and general generative AI ...
Master's degree or PhD in Engineering, Computer Science, or a related technical field. * 5 years of experience with LLM evaluation, machine learning algorithms and tools, and general generative AI ...
The ideal candidate possesses a solid understanding of machine learning, natural language understanding, modern LLM architectures, LLM evaluation & tooling, and a passion for pushing boundaries in ...
The ideal candidate possesses a solid understanding of machine learning, natural language understanding, modern LLM architectures, LLM evaluation & tooling, and a passion for pushing boundaries in ...
Edison, NJ · On-site
LLM Evaluation & Optimization * Python / C# / JavaScript Development * REST APIs & Enterprise Integration Patterns * Azure Cloud Architecture * AI Solution Architecture & Design * Responsible AI ...
Quick apply
Edison, NJ · On-site
LLM Evaluation & Optimization * Python / C# / JavaScript Development * REST APIs & Enterprise Integration Patterns * Azure Cloud Architecture * AI Solution Architecture & Design * Responsible AI ...
Manhattan, NY · Remote
$30 - $50/hr
Highlight STEM engineering impact, Python/data tooling, and experience with data labeling, RLHF, LLM evaluation, training data quality, and annotation guidelines compliance.
New
Manhattan, NY · Remote
$30 - $50/hr
Highlight STEM engineering impact, Python/data tooling, and experience with data labeling, RLHF, LLM evaluation, training data quality, and annotation guidelines compliance.
New
The ideal candidate possesses a solid understanding of machine learning, natural language understanding, modern LLM architectures, LLM evaluation & tooling, and a passion for pushing boundaries in ...
The ideal candidate possesses a solid understanding of machine learning, natural language understanding, modern LLM architectures, LLM evaluation & tooling, and a passion for pushing boundaries in ...
Manhattan, NY · Remote
$30 - $50/hr
Highlight STEM engineering impact, Python/data tooling, and experience with data labeling, RLHF, LLM evaluation, training data quality, and annotation guidelines compliance.
New
Manhattan, NY · Remote
$30 - $50/hr
Highlight STEM engineering impact, Python/data tooling, and experience with data labeling, RLHF, LLM evaluation, training data quality, and annotation guidelines compliance.
New
Implement LLM evaluation frameworks using LangSmith, Ragas, DeepEval, or custom evaluation pipelines to measure answer quality, groundedness, latency, and hallucination rates. * Apply LLMOps best ...
Implement LLM evaluation frameworks using LangSmith, Ragas, DeepEval, or custom evaluation pipelines to measure answer quality, groundedness, latency, and hallucination rates. * Apply LLMOps best ...
New York, NY · On-site
AI/LLM evaluation and LLM-as-a-Judge * Embeddings/vector search * Strong analytical and research skills * Experience with clinical, financial, or regulated domains is a plus.
New York, NY · On-site
AI/LLM evaluation and LLM-as-a-Judge * Embeddings/vector search * Strong analytical and research skills * Experience with clinical, financial, or regulated domains is a plus.
AI Evaluation Experience Previous experience with AI evaluation, RLHF, prompt engineering, AI annotation, or LLM evaluation is strongly preferred . Candidates with exceptional finance expertise who ...
AI Evaluation Experience Previous experience with AI evaluation, RLHF, prompt engineering, AI annotation, or LLM evaluation is strongly preferred . Candidates with exceptional finance expertise who ...
AI Evaluation Experience Previous experience with AI evaluation, RLHF, prompt engineering, AI annotation, or LLM evaluation is strongly preferred . Candidates with exceptional finance expertise who ...
AI Evaluation Experience Previous experience with AI evaluation, RLHF, prompt engineering, AI annotation, or LLM evaluation is strongly preferred . Candidates with exceptional finance expertise who ...
New York, NY · On-site
$107K - $137K/yr
Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency. * Build ...
New York, NY · On-site
$107K - $137K/yr
Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency. * Build ...
Own the diagnostic loop for LLM-based clinical services: take failure modes surfaced by clinical ... Own the evaluation roadmap and quality bar for the team's AI services: which metrics gate ...
New
Own the diagnostic loop for LLM-based clinical services: take failure modes surfaced by clinical ... Own the evaluation roadmap and quality bar for the team's AI services: which metrics gate ...
New
Implement LLM-based workflows including prompt engineering, evaluation, and RAG * Build knowledge retrieval pipelines to support IVR use cases (FAQs, troubleshooting, account queries) * Collaborate ...
Implement LLM-based workflows including prompt engineering, evaluation, and RAG * Build knowledge retrieval pipelines to support IVR use cases (FAQs, troubleshooting, account queries) * Collaborate ...
Lake Mary, FL · On-site
LLM Evaluation metrics (accuracy, faithfulness and relevance) * Prompt/version tracking * Logging and Tracking * Human in the loop (HITL) feedback loops * Deployment and MLOPS * Containerization ...
Quick apply
Lake Mary, FL · On-site
LLM Evaluation metrics (accuracy, faithfulness and relevance) * Prompt/version tracking * Logging and Tracking * Human in the loop (HITL) feedback loops * Deployment and MLOPS * Containerization ...
AI Evaluation Experience Previous experience with AI evaluation, RLHF, prompt engineering, AI annotation, or LLM evaluation is strongly preferred . Candidates with exceptional finance expertise who ...
AI Evaluation Experience Previous experience with AI evaluation, RLHF, prompt engineering, AI annotation, or LLM evaluation is strongly preferred . Candidates with exceptional finance expertise who ...
| Aspect | Llm Evaluation | Data Scientist |
|---|---|---|
| Required Credentials | Typically requires knowledge of machine learning, NLP, and AI concepts; often a degree in computer science or related fields | Requires degrees in computer science, statistics, or related fields; often includes certifications in data analysis or machine learning |
| Work Environment | Primarily research and testing environments, focusing on AI model assessment | Data analysis, modeling, and visualization in various industries like finance, healthcare, or tech |
| Employer & Industry Usage | Used by AI research labs, tech companies, and organizations developing NLP models | Used across industries for data analysis, predictive modeling, and business insights |
While both roles involve working with data and machine learning, Llm Evaluation focuses on assessing large language models' performance, whereas Data Scientists develop and implement data-driven solutions across various sectors.
Cities with the most Llm Evaluation job openings:
States with the most job openings for Llm Evaluation jobs include:
The top searched job categories for Llm Evaluation jobs are:

Sourced by ZipRecruiter
It services
51 - 200 Employees
Fremont, CA, US
2012