325 citations · 634 across the 31 of their papers we have counts for
1 paper · 2 filters
Eddie Yang, Dashun Wang
Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep e…