1 paper · 1 filter
Haoze Du, Richard Li, Edward Gehringer
Evaluating the performance of Large Language Models (LLMs) is a critical yet challenging task, particularly when aiming to avoid subjective assessments. This paper proposes a frame…