1 paper
Yongjie Wang, Xinyue Zhang, Kunhong Yao +4
Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference. Such age…