1 paper
Boyan Li, Bingsen Chen, Chenghao Yang +3
Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have emerged as two dominant methods for post-training reasoning LLMs. Prior work uses OPD's…