collaborators

6 papers

cs.LG2026

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay

Xiaolong Li, Zhangchen Zhou, Zhi-Qin John Xu

Most explanations of training instability focus on \emph{learning-rate criticality}, typically characterized by the Edge of Stability, beyond which optimization becomes unstable. W…

cs.LG2026

Towards Understanding Adam Convergence on Highly Degenerate Polynomials

Zhiwei Bai, Jiajie Zhao, Zhangchen Zhou +2

Adam is a widely used optimization algorithm in deep learning, yet the specific class of objective functions where it exhibits inherent advantages remains underexplored. Unlike pri…

cs.LG2026

Adaptive Preconditioners Trigger Loss Spikes in Adam

Zhiwei Bai, Zhangchen Zhou, Jiajie Zhao +6

Loss spikes commonly emerge during neural network training with the Adam optimizer across diverse architectures and scales, yet their underlying mechanism remains elusive. While pr…

cs.LG2026

An overview of condensation phenomenon in deep learning

Zhi-Qin John Xu, Yaoyu Zhang, Zhangchen Zhou

In this paper, we provide an overview of a common phenomenon, condensation, observed during the nonlinear training of neural networks: During the nonlinear training of neural netwo…

cs.AI2025

Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism

Zhiwei Wang, Yunji Wang, Zhongwang Zhang +7

Large language models have consistently struggled with complex reasoning tasks, such as mathematical problem-solving. Investigating the internal reasoning mechanisms of these model…

cs.LG2025

Scalable Complexity Control Facilitates Reasoning Ability of LLMs

Liangkai Hang, Junjie Yao, Zhiwei Bai +17

The reasoning ability of large language models (LLMs) has been rapidly advancing in recent years, attracting interest in more fundamental approaches that can reliably enhance their…