1 paper
Longyuan Zhu, Hairan Hua, Linlin Miao +1
Large Language Models (LLMs) are advancing rapidly, yet the benchmarks used to measure this progress are becoming increasingly unreliable. Score inflation and selective reporting h…