12 citations · 20 across the 10 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Escaping the KL Agreement Trap in On-Policy Distillation
Haoran Xin, Anhao Zhao, Ying Sun +3
On-policy distillation (OPD) provides dense token-level supervision by asking a teacher to score student-generated rollouts. However, when the student drifts into an unrecoverable…
cs.LG2026
GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
He Zhang, Ying Sun, Hui Xiong
Flow-matching policies hold great promise for reinforcement learning (RL) by capturing complex, multi-modal action distributions. However, their practical application is often hind…