3 papers
cs.LG2026
Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay
Xiaolong Li, Zhangchen Zhou, Zhi-Qin John Xu
Most explanations of training instability focus on \emph{learning-rate criticality}, typically characterized by the Edge of Stability, beyond which optimization becomes unstable. W…
cs.LG2025
Scalable Complexity Control Facilitates Reasoning Ability of LLMs
Liangkai Hang, Junjie Yao, Zhiwei Bai +17
The reasoning ability of large language models (LLMs) has been rapidly advancing in recent years, attracting interest in more fundamental approaches that can reliably enhance their…
cs.LG2025
An overview of condensation phenomenon in deep learning
Zhi-Qin John Xu, Yaoyu Zhang, Zhangchen Zhou
In this paper, we provide an overview of a common phenomenon, condensation, observed during the nonlinear training of neural networks: During the nonlinear training of neural netwo…