Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory
Runxi Cheng, Yuchen Guan, Yongxian Wei +7
Scaling conditional memory offers a promising way to increase language-model capacity, but existing methods such as Engram learn large memory tables from scratch during pre-trainin…
cs.CL2026
Spectral Characterization and Mitigation of Sequential Knowledge Editing Collapse
Chi Zhang, Mengqi Zhang, Xiaotian Ye +5
Sequential knowledge editing in large language models often causes catastrophic collapse of the model's general abilities, especially for parameter-modifying methods. Existing appr…
cs.CL2025
Mixture of Neuron Experts
Runxi Cheng, Yuchen Guan, Yucheng Ding +6
In this work, we first explore whether the parameters activated by the MoE layer remain highly sparse at inference. We perform a sparsification study on several representative MoE…