1 paper
Zongqi Wang, Tianle Gu, Chen Gong +3
Evaluating Large Language Models (LLMs) has become increasingly important, with automatic evaluation benchmarks gaining prominence as alternatives to human evaluation. While existi…