4 papers
Focus and Dilution: The Multi-stage Learning Process of Attention
Zheng-An Chen, Pengxiao Lin, Zhi-Qin John Xu +1
Transformer-based models have achieved remarkable success across a wide range of domains, yet our understanding of their training dynamics remains limited. In this work, we identif…
From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics
Zheng-An Chen, Tao Luo
Although transformer-based models have shown exceptional empirical performance, the fundamental principles governing their training dynamics are inadequately characterized beyond c…
Semi-Discrete in Time Method for Time-Dependent Equations by Random Neural Basis
Guihong Wang, Zheng-An Chen, Tao Luo
Neural network-based solvers for partial differential equations (PDEs) have attracted considerable attention, yet they often face challenges in accuracy and computational efficienc…
On Multi-Stage Loss Dynamics in Neural Networks: Mechanisms of Plateau and Descent Stages
Zheng-An Chen, Tao Luo, GuiHong Wang
The multi-stage phenomenon in the training loss curves of neural networks has been widely observed, reflecting the non-linearity and complexity inherent in the training process. In…