paper

Natasha 2: Faster Non-Convex Optimization Than SGD

arXiv:1708.08694

Abstract

We design a stochastic algorithm to train any smooth neural network to -approximate local minima, using backpropagations. The best result was essentially by SGD. More broadly, it finds -approximate local minima of any smooth nonconvex function in rate , with only oracle access to stochastic gradients.

V2 and V3 polished writing; V4 was a deep revision and simplified proofs

Cited by in corpus (14)