activity
20152021
most citedLearning One-hidden-layer Neural Networks with Landscape Design

112 citations · 288 across the 11 of their papers we have counts for

collaborators
Showing stat.MLShow all

9 papers · 1 filter

stat.ML2020

Beyond Lazy Training for Over-parameterized Tensor Decomposition

Xiang Wang, Chenwei Wu, Jason D. Lee +2

Over-parametrization is an important technique in training neural networks. In both theory and practice, training a larger network allows the optimization algorithm to avoid bad lo…

stat.ML2020

Modeling from Features: a Mean-field Framework for Over-parameterized Deep Neural Networks

Cong Fang, Jason D. Lee, Pengkun Yang +1

This paper proposes a new mean-field framework for over-parameterized deep neural networks (DNNs), which can be used to analyze neural network training. In this framework, a DNN is…

stat.ML201923 cited

Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models

Mor Shpigel Nacson, Suriya Gunasekar, Jason D. Lee +2

With an eye toward understanding complexity control in deep learning, we study how infinitesimal regularization or gradient descent optimization lead to margin maximizing solutions…

stat.ML2018

Regularization Matters: Generalization and Optimization of Neural Nets v.s. their Induced Kernel

Colin Wei, Jason D. Lee, Qiang Liu +1

Recent works have shown that on sufficiently over-parametrized neural nets, gradient descent with relatively large initialization optimizes a prediction function in the RKHS of the…

stat.ML2018

Adding One Neuron Can Eliminate All Bad Local Minima

Shiyu Liang, Ruoyu Sun, Jason D. Lee +1

One of the main difficulties in analyzing neural networks is the non-convexity of the loss function which may have many bad local minima. In this paper, we study the landscape of n…

stat.ML2018

Convergence of Gradient Descent on Separable Data

Mor Shpigel Nacson, Jason D. Lee, Suriya Gunasekar +3

We provide a detailed study on the implicit bias of gradient descent when optimizing loss functions with strictly monotone tails, such as the logistic loss, over separable datasets…