1 paper
Junhao Chen, Jingbo Sun, Xiang Li +4
As large language models (LLMs) advance across diverse tasks, the need for comprehensive evaluation beyond single metrics becomes increasingly important. To fully assess LLM intell…