1 paper
Yulong Chen, Yang Liu, Jianhao Yan +6
The impressive performance of Large Language Models (LLMs) has consistently surpassed numerous human-designed benchmarks, presenting new challenges in assessing the shortcomings of…