1 paper · 1 filter
Jinchang Zhu, Jindong Li, Chengyu Zou +4
Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking, each target token's effecti…