2 papers
cs.LG2026
The Stability of Singular Distribution: A Spectral Perspective on the Two-Phase Dynamics of Language Model Pre-training
Hongtao Zhang, Wenjie Zhou, Chenxi Jia +2
Large language model pre-training typically exhibits a two-phase trajectory: a fast initial loss drop followed by a prolonged slow improvement. We identify an underlying spectral p…
cs.LG2026
Extra-Merge: Tracing the Rank-1 Subspace of Model Merging in Language Model Pre-Training
Wenjie Zhou, Bohan Wang, Hongtao Zhang +3
Model merging has emerged as a lightweight paradigm for enhancing Large Language Models (LLMs), yet its underlying mechanisms remain poorly understood. In this work, we analyze lat…