1 paper
Wenze Lin, Jiale Zhao, Xitai Jiang +5
Reinforcement Learning with Verifiable Rewards (RLVR) and on-policy distillation (OPD) have become two widely adopted paradigms for post-training large language models. However, RL…