3 papers
cs.LG2026
Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining
Zihan Liu, Ruiheng Zheng, Shaobo Zhang +4
We uncover ELR collapse in language model pretraining: learning rate (LR) and parameter norm govern loss dynamics primarily through their ratio, the effective learning rate (ELR).…
cs.LG2025
The Sharpness Disparity Principle in Transformers for Accelerating Language Model Pre-Training
Jinbo Wang, Mingze Wang, Zhanpeng Zhou +3
Transformers consist of diverse building blocks, such as embedding layers, normalization layers, self-attention mechanisms, and point-wise feedforward networks. Thus, understanding…
cs.LG2024
How Transformers Get Rich: Approximation and Dynamics Analysis
Mingze Wang, Ruoxi Yu, Weinan E +1
Transformers have demonstrated exceptional in-context learning capabilities, yet the theoretical understanding of the underlying mechanisms remains limited. A recent work (Elhage e…