1 paper · 1 filter
Yanyu Chen, Jiyue Jiang, Jiahong Liu +3
The evaluation of Deep Research Agents is a critical challenge, as conventional outcome-based metrics fail to capture the nuances of their complex reasoning. Current evaluation fac…