2 papers
cs.CL2025
Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework
Zishan Xu, Shuyi Xie, Qingsong Lv +4
With the widespread application of Large Language Models (LLMs) in various tasks, the mainstream LLM platforms generate massive user-model interactions daily. In order to efficient…
cs.CL2024
IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation
Fan Lin, Shuyi Xie, Yong Dai +7
As Large Language Models (LLMs) grow increasingly adept at managing complex tasks, the evaluation set must keep pace with these advancements to ensure it remains sufficiently discr…