3 papers
stat.ML2026
Learning Perturbations to Extrapolate Your LLM
Zetai Cen, Chenfei Gu, Jin Zhu +3
Recent advancements in large language models demonstrate that injecting perturbations can substantially enhance extrapolation performance. However, current approaches often rely on…
cs.CL2026
Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
Jinchang Zhu, Jindong Li, Chengyu Zou +4
Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking, each target token's effecti…
cs.CL2026
Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
Jinchang Zhu, Jindong Li, Yuwen Hao +3
A causal-decoder block is hierarchical: lower layers build the residual basis that upper layers attend over. We identify a failure mode in GPT pretraining: upper layers commit to s…