2 papers
cs.LG2026
CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts
Naibin Gu, Qingyi Si, Chenxu Yang +5
On-policy distillation (OPD) provides dense token-level supervision on student-generated trajectories. Although OPD performs strongly when teacher and student belong to the same mo…
cs.LG2026
Near-Future Policy Optimization
Chuanyu Qin, Chenxu Yang, Qingyi Si +6
Reinforcement learning with verifiable rewards (RLVR) has become a core post-training recipe. Introducing suitable off-policy trajectories into on-policy exploration accelerates RL…