activity
20172023
most citedFine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks

256 citations · 627 across the 29 of their papers we have counts for

collaborators
Showing 2018Show all

11 papers · 1 filter

cs.LG2018

Gradient Descent Finds Global Minima of Deep Neural Networks

Simon S. Du, Jason D. Lee, Haochuan Li +2

Gradient descent finds a global minimum in training deep neural networks despite the objective function being non-convex. The current paper proves gradient descent achieves zero tr…

cs.LG2018

Gradient Descent Provably Optimizes Over-parameterized Neural Networks

Simon S. Du, Xiyu Zhai, Barnabas Poczos +1

One of the mysteries in the success of neural networks is randomly initialized first order methods like gradient descent can achieve zero training loss even though the objective fu…

cs.AI2018

Discrete-Continuous Mixtures in Probabilistic Programming: Generalized Semantics and Inference Algorithms

Yi Wu, Siddharth Srivastava, Nicholas Hay +2

Despite the recent successes of probabilistic programming languages (PPLs) in AI applications, PPLs offer only limited support for random variables whose distributions combine disc…

cs.LG2018

Algorithmic Regularization in Learning Deep Homogeneous Models: Layers are Automatically Balanced

Simon S. Du, Wei Hu, Jason D. Lee

We study the implicit regularization imposed by gradient descent for learning multi-layer homogeneous functions including feed-forward fully connected and convolutional deep neural…

cs.LG2018

Improved Learning of One-hidden-layer Convolutional Neural Networks with Overlaps

Simon S. Du, Surbhi Goel

We propose a new algorithm to learn a one-hidden-layer convolutional neural network where both the convolutional weights and the outputs weights are parameters to be learned. Our a…

math.ST2018

Robust Nonparametric Regression under Huber's -contamination Model

Simon S. Du, Yining Wang, Sivaraman Balakrishnan +2

We consider the non-parametric regression problem under Huber's -contamination model, in which an fraction of observations are subject to arbitrary adversarial noise. We fir…