1 paper
Jiaxuan Wang, Xuan Ouyang, Zhiyu Chen +4
On-policy self-distillation (self-OPD) densifies reinforcement learning with verifiable rewards (RLVR) by letting a policy teach itself under privileged context. We find that when…