1 paper
Yiran Zhang, Mo Wang, Xiaoyang Li +3
Despite impressive advances in large language models (LLMs), existing benchmarks often focus on single-turn or single-step tasks, failing to capture the kind of iterative reasoning…