1 paper · 1 filter
Jifan Yu, Xiaozhi Wang, Shangqing Tu +32
The unprecedented performance of large language models (LLMs) necessitates improvements in evaluations. Rather than merely exploring the breadth of LLM abilities, we believe meticu…