21 citations · 73 across the 40 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
Yang Sun, Lichao Ma, Houyuan Qin +5
On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolation methods such as ExOPD amplify the tea…
cs.LG2026
Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update
Daocheng Fu, Rong Wu, Yu Yang +7
Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward behaviors from policy exploration. While on…