7 papers
Escaping the KL Agreement Trap in On-Policy Distillation
Haoran Xin, Anhao Zhao, Ying Sun +3
On-policy distillation (OPD) provides dense token-level supervision by asking a teacher to score student-generated rollouts. However, when the student drifts into an unrecoverable…
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
Mingxu Zhang, Yuhan Li, Lujundong Li +3
Continual learning enables large language models to adapt to evolving tasks without retraining from scratch, yet catastrophic forgetting remains a central obstacle. Among continual…
Discrete Preference Learning for Personalized Multimodal Generation
Yuting Zhang, Ying Sun, Dazhong Shen +6
The emergence of generative models enables the creation of texts and images tailored to users' preferences. Existing personalized generative models have two critical limitations: l…
GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies
He Zhang, Ying Sun, Hui Xiong
Flow-matching policies hold great promise for reinforcement learning (RL) by capturing complex, multi-modal action distributions. However, their practical application is often hind…
Efficient Skill Discovery via Regret-Aware Optimization
He Zhang, Ming Zhou, Shaopeng Zhai +2
Unsupervised skill discovery aims to learn diverse and distinguishable behaviors in open-ended reinforcement learning. For existing methods, they focus on improving diversity throu…
Improving Recommendation Fairness without Sensitive Attributes Using Multi-Persona LLMs
Haoran Xin, Ying Sun, Chao Wang +3
Despite the success of recommender systems in alleviating information overload, fairness issues have raised concerns in recent years, potentially leading to unequal treatment for c…