1 paper · 1 filter
Bin Lei, Yu Li, Prafulla Kumar Choubey +7
Tree-structured rollouts give critic-free reinforcement learning with verifiable rewards (RLVR) step-level credit: fork a chain at an intermediate point, and sibling outcome differ…