5 papers
Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks
Wenquan Ma, Yang Sui, Jiaye Teng +3
Algorithmic stability is among the most potent techniques in generalization analysis. However, its derivation usually requires a stepsize under non-convex…
BSFA: Leveraging the Subspace Dichotomy to Accelerate Neural Network Training
Wenjie Zhou, Bohan Wang, Wei Chen +1
Recent studies \citep{gur2018gradient,song2024does, wen2024understanding} highlight a fundamental dichotomy in deep learning optimization: Although parameter updates along the top…
Continuous-time Riemannian SGD and SVRG Flows on Wasserstein Probabilistic Space
Mingyang Yi, Bohan Wang
Recently, optimization on the Riemannian manifold have provided valuable insights to the optimization community. In this regard, extending these methods to to the Wasserstein space…
Solving Formal Math Problems by Decomposition and Iterative Reflection
Yichi Zhou, Jianqiu Zhao, Yongxin Zhang +14
General-purpose Large Language Models (LLMs) have achieved remarkable success in intelligence, performing comparably to human experts on complex reasoning tasks such as coding and…
AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-training
Huishuai Zhang, Bohan Wang, Luoxin Chen
We introduce AdamS, a simple yet effective alternative to Adam for large language model (LLM) pretraining and post-training. By leveraging a novel denominator, i.e., the root of we…