1 paper
Zheng Liu, Longxiang Zhang, Xintong Wang +8
LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on outcome-homogeneous groups wh…