3 papers
cs.LG2026
Denser Better: Limits of On-Policy Self-Distillation for Continual Post-Training
Meng Wang, Haohan Zhao, Wenzhuo Liu +7
Continual post-training enables foundation models to acquire new knowledge while preserving existing capabilities. Recent work suggests that on-policy learning can mitigate forgett…
cs.LG2026
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
Song Lai, Haohan Zhao, Rong Feng +9
Continual post-training (CPT) is a popular and effective technique for adapting foundation models like multimodal large language models to ever-evolving downstream tasks. While exi…
cs.AI2025
The Docking Game: Loop Self-Play for Fast, Dynamic, and Accurate Prediction of Flexible Protein-Ligand Binding
Youzhi Zhang, Yufei Li, Gaofeng Meng +2
Molecular docking is a crucial aspect of drug discovery, as it predicts the binding interactions between small-molecule ligands and protein pockets. However, current multi-task lea…