1 paper
Xiang Li, Yunshi Lan, Chao Yang
Recently, numerous new benchmarks have been established to evaluate the performance of large language models (LLMs) via either computing a holistic score or employing another LLM a…