Monitoring and evaluation tools like LangSmith, LLM deployment libraries like LiteLLM, LLM Guardrails, and more, CI/CD, model versioning, and model deployment best practices • Knowledge of or ...
Monitoring and evaluation tools like LangSmith, LLM deployment libraries like LiteLLM, LLM Guardrails, and more, CI/CD, model versioning, and model deployment best practices • Knowledge of or ...
... and LLM evaluation workflows • Familiarity with enterprise AI tools such as Claude, GPT-4/4o, ChatGPT Enterprise, Copilot, Glean, and Shelf • Ingest and process unstructured/structured data ...
... and LLM evaluation workflows • Familiarity with enterprise AI tools such as Claude, GPT-4/4o, ChatGPT Enterprise, Copilot, Glean, and Shelf • Ingest and process unstructured/structured data ...
AI/ML Engineer
Minneapolis, MN · On-site
AI/ML Engineering & MLOps Experience developing, evaluating, deploying, and monitoring ML/LLM solutions, including LLM evaluation metrics, hallucination reduction, safety guardrails, CI/CD, cloud ...
AI/ML Engineer
Minneapolis, MN · On-site
AI/ML Engineering & MLOps Experience developing, evaluating, deploying, and monitoring ML/LLM solutions, including LLM evaluation metrics, hallucination reduction, safety guardrails, CI/CD, cloud ...
IT - Technology Lead | OpenSystem | Python - OpenSystem
Charlotte, NC · On-site
$131K - $161K/yr
... LLM evaluation Familiarity with MLOps tools and practices, including Monitoring and evaluation tools like LangSmith LLM deployment libraries like LiteLLM, LLM Guardrails, and more CICD, model ...
IT - Technology Lead | OpenSystem | Python - OpenSystem
Charlotte, NC · On-site
$131K - $161K/yr
... LLM evaluation Familiarity with MLOps tools and practices, including Monitoring and evaluation tools like LangSmith LLM deployment libraries like LiteLLM, LLM Guardrails, and more CICD, model ...
LLM evaluation, guardrails, privacy, bias, and safety * Scalable, low-latency, multi-tenant distributed systems design * Observability, SLOs, incident response, capacity, and cost management
LLM evaluation, guardrails, privacy, bias, and safety * Scalable, low-latency, multi-tenant distributed systems design * Observability, SLOs, incident response, capacity, and cost management
LLM evaluation, guardrails, privacy, bias, and safety * Scalable, low-latency, multi-tenant distributed systems design * Observability, SLOs, incident response, capacity, and cost management
LLM evaluation, guardrails, privacy, bias, and safety * Scalable, low-latency, multi-tenant distributed systems design * Observability, SLOs, incident response, capacity, and cost management
Data Scientist, Onsite We have an immediate need for an experienced Data Scientist with strong expertise in Generative AI, LLM evaluation, conversational AI, statistical analysis, and machine ...
Data Scientist, Onsite We have an immediate need for an experienced Data Scientist with strong expertise in Generative AI, LLM evaluation, conversational AI, statistical analysis, and machine ...
Familiarity with LLM evaluation metrics and prompt tuning techniques * Experience validating models integrated with Epic or EHR apps
Quick apply
Familiarity with LLM evaluation metrics and prompt tuning techniques * Experience validating models integrated with Epic or EHR apps
LLM evaluation, guardrails, privacy, bias, and safety * Scalable, low-latency, multi-tenant distributed systems design * Observability, SLOs, incident response, capacity, and cost management
LLM evaluation, guardrails, privacy, bias, and safety * Scalable, low-latency, multi-tenant distributed systems design * Observability, SLOs, incident response, capacity, and cost management
Qualifications * 3+ years of hands-on experience in LLM / GenAI data evaluation. * Bachelor's Degree required * Ability to research unfamiliar topics using trusted sources and make well-supported ...
Qualifications * 3+ years of hands-on experience in LLM / GenAI data evaluation. * Bachelor's Degree required * Ability to research unfamiliar topics using trusted sources and make well-supported ...
LLM Model Response Evaluation
Maryland City, MD · On-site +1
Qualifications * 3+ years of hands-on experience in LLM / GenAI data evaluation. * Bachelor's Degree required * Ability to research unfamiliar topics using trusted sources and make well-supported ...
LLM Model Response Evaluation
Maryland City, MD · On-site +1
Qualifications * 3+ years of hands-on experience in LLM / GenAI data evaluation. * Bachelor's Degree required * Ability to research unfamiliar topics using trusted sources and make well-supported ...
LLM evaluation, guardrails, privacy, bias, and safety * Scalable, low-latency, multi-tenant distributed systems design * Observability, SLOs, incident response, capacity, and cost management
LLM evaluation, guardrails, privacy, bias, and safety * Scalable, low-latency, multi-tenant distributed systems design * Observability, SLOs, incident response, capacity, and cost management
Technical Solutions Architect, Evals & Fine-Tuning
$140K - $160K/yr
SFT data curation, preference data collection for RLHF/DPO, golden datasets, custom benchmarks, LLM-as-judge pipelines, human-in-the-loop evaluation, red teaming, and multimodal eval (text, image ...
Technical Solutions Architect, Evals & Fine-Tuning
$140K - $160K/yr
SFT data curation, preference data collection for RLHF/DPO, golden datasets, custom benchmarks, LLM-as-judge pipelines, human-in-the-loop evaluation, red teaming, and multimodal eval (text, image ...
AI Engineer
Manhattan, NY · On-site
... evaluation frameworks, including datasets and benchmarks, to ensure our LLM applications meet quality standards and perform reliably over time. • Provide mentorship to team members, sharing your ...
AI Engineer
Manhattan, NY · On-site
... evaluation frameworks, including datasets and benchmarks, to ensure our LLM applications meet quality standards and perform reliably over time. • Provide mentorship to team members, sharing your ...
LLM Ops Engineer
Denver, CO · Hybrid
$105K - $130K/yr
As an LLM Ops Engineer, you will create the secure, scalable, and reliable foundation that allows ... Lead the evolution of Litera's AI operations capabilities by evaluating emerging technologies and ...
LLM Ops Engineer
Denver, CO · Hybrid
$105K - $130K/yr
As an LLM Ops Engineer, you will create the secure, scalable, and reliable foundation that allows ... Lead the evolution of Litera's AI operations capabilities by evaluating emerging technologies and ...
LLM Ops Engineer
Denver, CO · On-site
$105K - $130K/yr
As an LLM Ops Engineer, you will create the secure, scalable, and reliable foundation that allows ... Lead the evolution of Litera's AI operations capabilities by evaluating emerging technologies and ...
LLM Ops Engineer
Denver, CO · On-site
$105K - $130K/yr
As an LLM Ops Engineer, you will create the secure, scalable, and reliable foundation that allows ... Lead the evolution of Litera's AI operations capabilities by evaluating emerging technologies and ...
Qualifications * 3+ years of hands-on experience in LLM / GenAI data evaluation. * Bachelor's Degree required * Ability to research unfamiliar topics using trusted sources and make well-supported ...
Qualifications * 3+ years of hands-on experience in LLM / GenAI data evaluation. * Bachelor's Degree required * Ability to research unfamiliar topics using trusted sources and make well-supported ...
Principal AI/ML Software Engineer
Houston, TX · On-site
$128K - $172K/yr
End-to-end experience with ML pipelines, model versioning, feature stores, drift detection, CI/CD for ML, and Docker containerization • LLM Evaluation: Experience with evaluation frameworks (RAGAS ...
Principal AI/ML Software Engineer
Houston, TX · On-site
$128K - $172K/yr
End-to-end experience with ML pipelines, model versioning, feature stores, drift detection, CI/CD for ML, and Docker containerization • LLM Evaluation: Experience with evaluation frameworks (RAGAS ...
Senior/Staff FDE - Synthetic Data Generation
New York, NY · On-site
$116K - $157K/yr
Develop LLM- and ML-assisted workflows to generate high-quality training and evaluation datasets across targeted behaviors, domains, and edge cases * Build automated evaluators, quality checks, and ...
Senior/Staff FDE - Synthetic Data Generation
New York, NY · On-site
$116K - $157K/yr
Develop LLM- and ML-assisted workflows to generate high-quality training and evaluation datasets across targeted behaviors, domains, and edge cases * Build automated evaluators, quality checks, and ...
AI/ML Software Engineer
$55 - $60/hr
Implement LLM evaluation, observability, guardrails, content filtering, and responsible AI practices. * Troubleshoot, optimize, and support production AI applications. Required Qualifications * 7+ ...
Quick apply
AI/ML Software Engineer
$55 - $60/hr
Implement LLM evaluation, observability, guardrails, content filtering, and responsible AI practices. * Troubleshoot, optimize, and support production AI applications. Required Qualifications * 7+ ...
Llm Evaluation information
What is the difference between Llm Evaluation vs Data Scientist?
| Aspect | Llm Evaluation | Data Scientist |
|---|---|---|
| Required Credentials | Typically requires knowledge of machine learning, NLP, and AI concepts; often a degree in computer science or related fields | Requires degrees in computer science, statistics, or related fields; often includes certifications in data analysis or machine learning |
| Work Environment | Primarily research and testing environments, focusing on AI model assessment | Data analysis, modeling, and visualization in various industries like finance, healthcare, or tech |
| Employer & Industry Usage | Used by AI research labs, tech companies, and organizations developing NLP models | Used across industries for data analysis, predictive modeling, and business insights |
While both roles involve working with data and machine learning, Llm Evaluation focuses on assessing large language models' performance, whereas Data Scientists develop and implement data-driven solutions across various sectors.
What cities are hiring for Llm Evaluation jobs?
Cities with the most Llm Evaluation job openings:
What states have the most Llm Evaluation jobs?
States with the most job openings for Llm Evaluation jobs include:
What job categories do people searching Llm Evaluation jobs look for?
The top searched job categories for Llm Evaluation jobs are:

Full-time
Re-posted 21 days ago
Job description
Info Way Solutions is a company seeking an AI/ML Engineer who is a native Spanish speaker. The role involves leveraging experience in AI/ML technologies, including Python and NLP, to contribute to various projects.
Qualifications:
Required:
• 6+ years overall experience, including 4+ years in AI/ML
• Proficiency in Python, ML, NLP, and text data analysis
• Hands-on experience with API creation, RAG, Vector databases, LLM prompt engineering, LLM evaluation
• Familiarity with MLOps tools and practices, including: Monitoring and evaluation tools like LangSmith, LLM deployment libraries like LiteLLM, LLM Guardrails, and more, CI/CD, model versioning, and model deployment best practices
• Knowledge of or willingness to learn Kubernetes and cloud services (AWS/GCP/Azure)
• Native Spanish Speaker
Company:
Founded and incorporated in 2012 , Info Way Solutions is an IT services and consulting company headquartered in Fremont , CA. Founded in 2012, the company is headquartered in Fremont, USA, with a team of 501-1000 employees. The company is currently Late Stage.
About Info Way Solutions
Sourced by ZipRecruiter
Industry
It services
Company size
51 - 200 Employees
Headquarters location
Fremont, CA, US
Year founded
2012