1 paper · 1 filter
Wentang Song, Jinqiang Li, Kele Huang +3
The versatility of Large Language Models (LLMs) in vertical domains has spurred the development of numerous specialized evaluation benchmarks. However, these benchmarks often suffe…