1 paper
Yuchen Liu, Yingjie Feng, Lixiong Qin +5
In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward methods typically rely on co…