1 paper · 1 filter
Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes…