1 paper
Alicia Sagae, Chia-Jung Lee, Sandeep Avula +2
Current methods for evaluating large language models (LLMs) typically focus on high-level tasks such as text generation, without targeting a particular AI application. This approac…