1 paper · 1 filter
Xu-Xiang Zhong, Chao Yi, Han-Jia Ye
With the development of Large Language Models (LLMs), numerous benchmarks have been proposed to measure and compare the capabilities of different LLMs. However, evaluating LLMs is…