3 papers
cs.LG2026
Pre-carved Niches: The Formation Dynamics of Modular Task Partitions in Early LLM Training
Guangqi Li, Yongxin Li
Large language models exhibit a modular internal organization that mirrors well-studied functional networks of the human brain, but how this organization forms during training is u…
cs.LG2026
Beyond What to Select: A Plug-and-play Oscillatory Data-Volume Scheduling for Efficient Model Training
Suorong Yang, Hanqi Zhu, Hai Gan +4
Data selection accelerates training by identifying representative training data while preserving model performance. However, existing methods mainly focus on designing sample-impor…
cs.CL2026
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
Xinle Deng, Ruobin Zhong, Hujin Peng +15
Memory is essential for enabling large language models to support long-horizon reasoning, yet existing memory systems remain unreliable and difficult to debug. Tracing memory's dyn…