1 paper
Xunyi Jiang, Dingyi Chang, Julian McAuley +1
The rapid evolution of large language models (LLMs) and the real world has outpaced the static nature of widely used evaluation benchmarks, raising concerns about their reliability…