activity
20242026
collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling

Yixiao Qian, Song Chen, Pengkai Wang +3

Modern language models are built primarily from Transformers, recurrent models, and their hybrid architectures. Transformers rely on token-level attention memories, while recurrent…

cs.LG2026

Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation

Yuanyi Wang, Su Lu, Yanggan Gu +6

On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uniformity of OPD signals by pr…

cs.LG2026

Discovering Physical Directions in Weight Space: Composing Neural PDE Experts

Pengkai Wang, Pengwei Liu, Yuanyi Wang +7

Recent advances in neural operators have made partial differential equation (PDE) surrogate modeling increasingly scalable and transferable through large-scale pretraining and in-c…

cs.LG2026

FeatCal: Feature Calibration for Post-Merging Models

Yanggan Gu, Shuo Cai, Zihao Wang +7

Model merging combines task experts into one model and avoids joint training, retraining, or deploying many expert models, but the merged model often still underperforms task exper…

cs.LG2026

Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training

Yuanyi Wang, Yifan Yang, Su Lu +9

Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential updates enable capability tran…