1 paper · 1 filter
Guoli Yin, Haoping Bai, Shuang Ma +21
Recent advances in large language models (LLMs) have increased the demand for comprehensive benchmarks to evaluate their capabilities as human-like agents. Existing benchmarks, whi…