306 citations · 332 across the 9 of their papers we have counts for
11 papers · 1 filter
When and Why Momentum Accelerates SGD:An Empirical Study
Jingwen Fu, Bohan Wang, Huishuai Zhang +3
Momentum has become a crucial component in deep learning optimizers, necessitating a comprehensive understanding of when and why it accelerates stochastic gradient descent (SGD). T…
Optimizing Information-theoretical Generalization Bounds via Anisotropic Noise in SGLD
Bohan Wang, Huishuai Zhang, Jieyu Zhang +3
Recently, the information-theoretical framework has been proven to be able to obtain non-vacuous generalization bounds for large models trained by Stochastic Gradient Langevin Dyna…
Regularized OFU: an Efficient UCB Estimator forNon-linear Contextual Bandit
Yichi Zhou, Shihong Song, Huishuai Zhang +3
Balancing exploration and exploitation (EE) is a fundamental problem in contex-tual bandit. One powerful principle for EE trade-off isOptimism in Face of Uncer-tainty(OFU), in whic…
R-Drop: Regularized Dropout for Neural Networks
Xiaobo Liang, Lijun Wu, Juntao Li +6
Dropout is a powerful and widely used technique to regularize the training of deep neural networks. In this paper, we introduce a simple regularization strategy upon dropout in mod…
Large Scale Private Learning via Low-rank Reparametrization
Da Yu, Huishuai Zhang, Wei Chen +2
We propose a reparametrization scheme to address the challenges of applying differentially private SGD on large neural networks, which are 1) the huge memory cost of storing indivi…
Incorporating NODE with Pre-trained Neural Differential Operator for Learning Dynamics
Shiqi Gong, Qi Meng, Yue Wang +4
Learning dynamics governed by differential equations is crucial for predicting and controlling the systems in science and engineering. Neural Ordinary Differential Equation (NODE),…