1 paper · 1 filter
Ruotian Peng, Yi Ren, Zhouliang Yu +2
We revisit exploration collapse in reinforcement learning with verifiable rewards (RLVR), from the perspective of the \emph{candidate distribution} for next-token prediction. We fo…