1 paper · 1 filter
Lele Liao, Qile Zhang, Ruofan Wu +1
Evaluating large language models (LLMs) on comprehensive benchmarks is a cornerstone of their development, yet it's often computationally and financially prohibitive. While Item Re…