256 citations · 627 across the 29 of their papers we have counts for
11 papers · 1 filter
Gradient Descent Finds Global Minima of Deep Neural Networks
Simon S. Du, Jason D. Lee, Haochuan Li +2
Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex. The current paper proves gradient descent achieves zero tr…
Gradient Descent Provably Optimizes Over-parameterized Neural Networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos +1
One of the mysteries in the success of neural networks is randomly initialized first order methods like gradient descent can achieve zero training loss even though the objective fu…
Discrete-Continuous Mixtures in Probabilistic Programming: Generalized Semantics and Inference Algorithms
Yi Wu, Siddharth Srivastava, Nicholas Hay +2
Despite the recent successes of probabilistic programming languages (PPLs) in AI applications, PPLs offer only limited support for random variables whose distributions combine disc…
Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically Balanced
Simon S. Du, Wei Hu, Jason D. Lee
We study the implicit regularization imposed by gradient descent for learning multi-layer homogeneous functions including feed-forward fully connected and convolutional deep neural…
Improved Learning of One-hidden-layer Convolutional Neural Networks with Overlaps
Simon S. Du, Surbhi Goel
We propose a new algorithm to learn a one-hidden-layer convolutional neural network where both the convolutional weights and the outputs weights are parameters to be learned. Our a…
Robust Nonparametric Regression under Huber's -contamination Model
Simon S. Du, Yining Wang, Sivaraman Balakrishnan +2
We consider the non-parametric regression problem under Huber's -contamination model, in which an fraction of observations are subject to arbitrary adversarial noise. We fir…