collaborators

7 papers

cs.LG2026

Escaping the KL Agreement Trap in On-Policy Distillation

Haoran Xin, Anhao Zhao, Ying Sun +3

On-policy distillation (OPD) provides dense token-level supervision by asking a teacher to score student-generated rollouts. However, when the student drifts into an unrecoverable…

cs.LG2026

SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models

Mingxu Zhang, Yuhan Li, Lujundong Li +3

Continual learning enables large language models to adapt to evolving tasks without retraining from scratch, yet catastrophic forgetting remains a central obstacle. Among continual…

cs.IR2026

Discrete Preference Learning for Personalized Multimodal Generation

Yuting Zhang, Ying Sun, Dazhong Shen +6

The emergence of generative models enables the creation of texts and images tailored to users' preferences. Existing personalized generative models have two critical limitations: l…

cs.LG2026

GoldenStart: Q-Guided Priors and Entropy Control for Distilling Flow Policies

He Zhang, Ying Sun, Hui Xiong

Flow-matching policies hold great promise for reinforcement learning (RL) by capturing complex, multi-modal action distributions. However, their practical application is often hind…

cs.LG2025

Efficient Skill Discovery via Regret-Aware Optimization

He Zhang, Ming Zhou, Shaopeng Zhai +2

Unsupervised skill discovery aims to learn diverse and distinguishable behaviors in open-ended reinforcement learning. For existing methods, they focus on improving diversity throu…

cs.IR2025

Improving Recommendation Fairness without Sensitive Attributes Using Multi-Persona LLMs

Haoran Xin, Ying Sun, Chao Wang +3

Despite the success of recommender systems in alleviating information overload, fairness issues have raised concerns in recent years, potentially leading to unequal treatment for c…