1 paper · 1 filter
Taolin Zhang, Hang Guo, Wang Lu +3
As large language models (LLMs) continue to scale up, their performance on various downstream tasks has significantly improved. However, evaluating their capabilities has become in…