1 paper · 1 filter
Jaehoon Kim, Dongha Lee
On-Policy Self-Distillation (OPSD) has recently emerged as an alternative to Reinforcement Learning with Verifiable Rewards (RLVR), promising higher accuracy and shorter responses…