1 paper · 1 filter
Boyan Li, Bingsen Chen, Chenghao Yang +3
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's…