11 papers
Hyperball May Not Be a Free Lunch
Yihao Xiao, Jialong Sun, Zitian Gao +5
For scale-invariant deep networks, Hyperball-style optimizers have shown strong performance in large-scale training by fixing the norms of matrix-valued parameters and normalizing…
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
Zixuan Gong, Shijia Li, Yong Liu +1
Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Counterfact dataset, the answers progress from…
Towards Understanding the Power and Limits of the Muon Optimizer: A River-Valley Perspective
Tianqi Shen, Jinji Yang, Runze Shi +3
Recently, Muon has gained substantial attention as an appealing alternative to Adam-like optimizers, with many works highlighting its advantages through spectral normalization and…
Questioning the Coverage-Length Metric in Conformal Prediction: When Shorter Intervals Are Not Better
Yizhou Min, Yizhou Lu, Lanqi Li +2
Conformal prediction(CP) has become a cornerstone of distribution-free uncertainty quantification, conventionally evaluated by its coverage and interval length. This work criticall…
Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
Huangyu Xu, Jingqin Yang, Qianqian Xu +1
Sparse optimization is a fundamental challenge in various practical applications. A popular approach to sparse optimization is regularization. However, it may encounter op…
Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks
Wenquan Ma, Yang Sui, Jiaye Teng +3
Algorithmic stability is among the most potent techniques in generalization analysis. However, its derivation usually requires a stepsize under non-convex…