1 paper
Hongli Zhou, Hui Huang, Ziqing Zhao +10
The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concern…