1

Llm Evaluation Jobs (NOW HIRING)

Own LLM evaluation processes and methods with a focus on generating benchmarks representative of real-world usage and safety vulnerabilities. * Generate high quality synthetic data, curate labels ...

NY · On-site

$150 - $200/hr

Your primary focus will be evaluating production LLM pipelines: designing accuracy metrics, measuring retrieval effectiveness, and making principled cost/quality tradeoffs for the prompts and models ...

LLM Platform Engineer

San Francisco, CA · On-site

$245K - $345K/yr

Create robust and scalable LLM evaluation frameworks to measure model performance, guide iteration, and prevent regression via CI/CD. * Deploy RAG systems and MCP servers to more effectively ground ...

Create robust and scalable LLM evaluation frameworks to measure model performance, guide iteration, and prevent regression via CI/CD. * Deploy RAG systems and MCP servers to more effectively ground ...

next page

Showing results 1-20

Llm Evaluation information

What is the difference between Llm Evaluation vs Data Scientist?

AspectLlm EvaluationData Scientist
Required CredentialsTypically requires knowledge of machine learning, NLP, and AI concepts; often a degree in computer science or related fieldsRequires degrees in computer science, statistics, or related fields; often includes certifications in data analysis or machine learning
Work EnvironmentPrimarily research and testing environments, focusing on AI model assessmentData analysis, modeling, and visualization in various industries like finance, healthcare, or tech
Employer & Industry UsageUsed by AI research labs, tech companies, and organizations developing NLP modelsUsed across industries for data analysis, predictive modeling, and business insights

While both roles involve working with data and machine learning, Llm Evaluation focuses on assessing large language models' performance, whereas Data Scientists develop and implement data-driven solutions across various sectors.

More about Llm Evaluation jobs

What cities are hiring for Llm Evaluation jobs?

Cities with the most Llm Evaluation job openings:

What states have the most Llm Evaluation jobs?

States with the most job openings for Llm Evaluation jobs include:

What job categories do people searching Llm Evaluation jobs look for?

The top searched job categories for Llm Evaluation jobs are:

Infographic showing various Llm Evaluation job openings in the United States as of August 2026, with employment types broken down into 2% As Needed, 81% Full Time, 15% Part Time, and 2% Contract. Highlights an 89% Physical, 2% Hybrid, and 9% Remote job distribution.

ML Engineer -- LLM Evaluation

Capitolis

San Francisco, CA • On-site

$125 - $150/hr

Other

Posted 6 days ago


Job description

At Dynamo AI, we believe that LLMs must be developed with safety, privacy, and real-world responsibility in mind. Our ML team comes from a culture of academic research driven to democratize AI advancements responsibly. By operating at the intersection of ML research and industry applications, our team empowers Fortune 500 companies’ adoption of frontier research for their next generation of LLM products.

Join us if you:

  • Wish to work on the premier platform for private and personalized LLMs. We provide the fastest end to end solution to deploy research in the real world with our fast-paced team of ML Ph.D.’s and builders, free of Big Tech / academic bureaucracy and constraints.
  • Are excited at the idea of democratizing state-of-the-art research on safe and responsible AI.
  • Are motivated to work at a 2023 CB Insights Top 100 AI Startup and see your impact on end customers in the timeframe of weeks not years.
  • Care about building a platform to empower fair, unbiased, and responsible development of LLMs and don’t accept the status quo of sacrificing user privacy for the sake of ML advancement.
Responsibilities
  • Own LLM evaluation processes and methods with a focus on generating benchmarks representative of real-world usage and safety vulnerabilities.
  • Generate high quality synthetic data, curate labels, and conduct rigorous benchmarking.
  • Deliver robust, scalable, and reproducible production code.
  • Push the envelope by developing methods for benchmarking that revamps how we assess the best LLMs for harmlessness and helpfulness. Your research will directly empower our customers to more feasibly deploy safe and responsible LLMs.
  • Co-author papers, patents, and presentations with our research team by integrating other members’ work with your vertical.
Qualifications
  • Domain knowledge in LLM evaluation and data curation techniques.
  • Extensive experience in designing and implementing LLM benchmarking, extending previous methods. Comfortability with leading end-to-end projects.
  • Adaptability and flexibility. In both the academic and startup world, a new finding in the community may necessitate an abrupt shift in focus. You must be able to learn, implement, and extend state-of-the-art research.
  • Preferred: past research or projects in benchmarking LLMs.

Dynamo AI is committed to maintaining compliance with all applicable local and state laws regarding job listings and salary transparency. This includes adhering to specific regulations that mandate the disclosure of salary ranges in job postings or upon request during the hiring process. We strive to ensure our practices promote fairness, equity, and transparency for all candidates.

Salary for this position may vary based on several factors, including the candidate's experience, expertise, and the geographic location of the role. Compensation is determined to ensure competitiveness and equity, reflecting the cost of living in different regions and the specific skills and qualifications of the candidate.

#J-18808-Ljbffr