2 papers
cs.LG2018
An Alternative View: When Does SGD Escape Local Minima?
Robert Kleinberg, Yuanzhi Li, Yang Yuan
Stochastic gradient descent (SGD) is widely used in machine learning. Although being commonly viewed as a fast but not accurate version of gradient descent (GD), it always finds be…
cs.LG2017
Convergence Analysis of Two-layer Neural Networks with ReLU Activation
Yuanzhi Li, Yang Yuan
In recent years, stochastic gradient descent (SGD) based techniques has become the standard tools for training neural networks. However, formal theoretical understanding of why SGD…