most citedSGD Converges to Global Minimum in Deep Learning via Star-convex Path

22 citations · 67 across the 5 of their papers we have counts for

collaborators

5 papers

cs.LG202017 cited

Reanalysis of Variance Reduced Temporal Difference Learning

Tengyu Xu, Zhe Wang, Yi Zhou +1

Temporal difference (TD) learning is a popular algorithm for policy evaluation in reinforcement learning, but the vanilla TD can substantially suffer from the inherent optimization…

cs.LG201916 cited

Improved Zeroth-Order Variance Reduced Algorithms and Analysis for Nonconvex Optimization

Kaiyi Ji, Zhe Wang, Yi Zhou +1

Two types of zeroth-order stochastic algorithms have recently been designed for nonconvex optimization respectively based on the first-order techniques SVRG and SARAH/SPIDER. This…

stat.ML20191 cited

Distributed SGD Generalizes Well Under Asynchrony

Jayanth Regatti, Gaurav Tendolkar, Yi Zhou +2

The performance of fully synchronized distributed systems has faced a bottleneck due to the big data trend, under which asynchronous distributed systems are becoming a major popula…

math.OC201911 cited

Momentum Schemes with Stochastic Variance Reduction for Nonconvex Composite Optimization

Yi Zhou, Zhe Wang, Kaiyi Ji +2

Two new stochastic variance-reduced algorithms named SARAH and SPIDER have been recently proposed, and SPIDER has been shown to achieve a near-optimal gradient oracle complexity fo…

cs.LG201922 cited

SGD Converges to Global Minimum in Deep Learning via Star-convex Path

Yi Zhou, Junjie Yang, Huishuai Zhang +2

Stochastic gradient descent (SGD) has been found to be surprisingly effective in training a variety of deep neural networks. However, there is still a lack of understanding on how…