387 citations · 451 across the 9 of their papers we have counts for
1 paper · 1 filter
Qi Jia, Haodong Zhao, Dun Pei +7
Benchmarking large language models (LLMs) and agents in multi-turn interactive scenarios is essential for understanding their practical capabilities. However, existing evaluation p…