1 paper
Fan Lin, Shuyi Xie, Yong Dai +7
As Large Language Models (LLMs) grow increasingly adept at managing complex tasks, the evaluation set must keep pace with these advancements to ensure it remains sufficiently discr…