112 citations · 295 across the 15 of their papers we have counts for
4 papers · 2 filters
Gradient Descent Finds Global Minima of Deep Neural Networks
Simon S. Du, Jason D. Lee, Haochuan Li +2
Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex. The current paper proves gradient descent achieves zero tr…
Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically Balanced
Simon S. Du, Wei Hu, Jason D. Lee
We study the implicit regularization imposed by gradient descent for learning multi-layer homogeneous functions including feed-forward fully connected and convolutional deep neural…
On the Power of Over-parametrization in Neural Networks with Quadratic Activation
Simon S. Du, Jason D. Lee
We provide new theoretical insights on why over-parametrization is effective in learning neural networks. For a hidden node shallow network with quadratic activation and tr…
On the Convergence and Robustness of Training GANs with Regularized Optimal Transport
Maziar Sanjabi, Jimmy Ba, Meisam Razaviyayn +1
Generative Adversarial Networks (GANs) are one of the most practical methods for learning data distributions. A popular GAN formulation is based on the use of Wasserstein distance…