5 papers · 1 filter
DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling
Yixiao Qian, Song Chen, Pengkai Wang +3
Modern language models are built primarily from Transformers, recurrent models, and their hybrid architectures. Transformers rely on token-level attention memories, while recurrent…
Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
Yuanyi Wang, Su Lu, Yanggan Gu +6
On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uniformity of OPD signals by pr…
Discovering Physical Directions in Weight Space: Composing Neural PDE Experts
Pengkai Wang, Pengwei Liu, Yuanyi Wang +7
Recent advances in neural operators have made partial differential equation (PDE) surrogate modeling increasingly scalable and transferable through large-scale pretraining and in-c…
FeatCal: Feature Calibration for Post-Merging Models
Yanggan Gu, Shuo Cai, Zihao Wang +7
Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often still underperforms task exper…
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
Yuanyi Wang, Yifan Yang, Su Lu +9
Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential updates enable capability tran…