306 citations · 338 across the 8 of their papers we have counts for
14 papers
Optimizing Information-theoretical Generalization Bounds via Anisotropic Noise in SGLD
Bohan Wang, Huishuai Zhang, Jieyu Zhang +3
Recently, the information-theoretical framework has been proven to be able to obtain non-vacuous generalization bounds for large models trained by Stochastic Gradient Langevin Dyna…
R-Drop: Regularized Dropout for Neural Networks
Xiaobo Liang, Lijun Wu, Juntao Li +6
Dropout is a powerful and widely used technique to regularize the training of deep neural networks. In this paper, we introduce a simple regularization strategy upon dropout in mod…
Incorporating NODE with Pre-trained Neural Differential Operator for Learning Dynamics
Shiqi Gong, Qi Meng, Yue Wang +4
Learning dynamics governed by differential equations is crucial for predicting and controlling the systems in science and engineering. Neural Ordinary Differential Equation (NODE),…
UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost
Zhen Wu, Lijun Wu, Qi Meng +5
Transformer architecture achieves great success in abundant natural language processing tasks. The over-parameterization of the Transformer model has motivated plenty of works to a…
Constructing Basis Path Set by Eliminating Path Dependency
Juanping Zhu, Qi Meng, Wei Chen +2
The way the basis path set works in neural network remains mysterious, and the generalization of newly appeared G-SGD algorithm to more practical network is hindered. The Basis Pat…
Dynamic of Stochastic Gradient Descent with State-Dependent Noise
Qi Meng, Shiqi Gong, Wei Chen +2
Stochastic gradient descent (SGD) and its variants are mainstream methods to train deep neural networks. Since neural networks are non-convex, more and more works study the dynamic…