4 papers
The Implicit Bias of Steepest Descent with Mini-batch Stochastic Gradient
Jichu Li, Xuan Tang, Difan Zou
A variety of widely used optimization methods like SignSGD and Muon can be interpreted as instances of steepest descent under different norm-induced geometries. In this work, we st…
A Mechanism Study of Delayed Loss Spikes in Batch-Normalized Linear Models
Peifeng Gao, Wenyi Fang, Yang Zheng +1
Delayed loss spikes have been reported in neural-network training, but existing theory mainly explains earlier non-monotone behavior caused by overly large fixed learning rates. We…
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
Xuan Tang, Jichu Li, Difan Zou
The rapid scaling of large language models (LLMs) has made low-precision training essential for reducing memory, improving efficiency, and enabling larger models and datasets. Exis…
Scaling Laws for Precision in High-Dimensional Linear Regression
Dechen Zhang, Xuan Tang, Yingyu Liang +1
Low-precision training is critical for optimizing the trade-off between model quality and training costs, necessitating the joint allocation of model size, dataset size, and numeri…