activity
20172021
most citedConvergence Analysis of Distributed Stochastic Gradient Descent with Shuffling

18 citations · 41 across the 7 of their papers we have counts for

collaborators

12 papers

cs.LG2021

Incorporating NODE with Pre-trained Neural Differential Operator for Learning Dynamics

Shiqi Gong, Qi Meng, Yue Wang +4

Learning dynamics governed by differential equations is crucial for predicting and controlling the systems in science and engineering. Neural Ordinary Differential Equation (NODE),…

cs.LG20218 cited

Improved OOD Generalization via Adversarial Training and Pre-training

Mingyang Yi, Lu Hou, Jiacheng Sun +4

Recently, learning a model that generalizes well on out-of-distribution (OOD) data has attracted great attention in the machine learning community. In this paper, after defining OO…

cs.LG20212 cited

Reweighting Augmented Samples by Minimizing the Maximal Expected Loss

Mingyang Yi, Lu Hou, Lifeng Shang +3

Data augmentation is an effective technique to improve the generalization of deep neural networks. However, previous data augmentation methods usually treat the augmented samples e…

cs.LG2021

BN-invariant sharpness regularizes the training model to better generalization

Mingyang Yi, Huishuai Zhang, Wei Chen +2

It is arguably believed that flatter minima can generalize better. However, it has been pointed out that the usual definitions of sharpness, which consider either the maxima or the…

cs.LG2020

Dynamic of Stochastic Gradient Descent with State-Dependent Noise

Qi Meng, Shiqi Gong, Wei Chen +2

Stochastic gradient descent (SGD) and its variants are mainstream methods to train deep neural networks. Since neural networks are non-convex, more and more works study the dynamic…

cs.LG20192 cited

Interpreting Basis Path Set in Neural Networks

Juanping Zhu, Qi Meng, Wei Chen +1

Based on basis path set, G-SGD algorithm significantly outperforms conventional SGD algorithm in optimizing neural networks. However, how the inner mechanism of basis paths work re…