DTN: A Learning Rate Scheme with Convergence Rate of for SGD
arXiv:1901.07634
Abstract
This paper has some inconsistent results, i.e., we made some failed claims because we did some mistakes for using the test criterion for a series. Precisely, our claims on the convergence rate of of SGD presented in Theorem 1, Corollary 1, Theorem 2 and Corollary 2 are wrongly derived because they are based on Lemma 5. In Lemma 5, we do not correctly use the test criterion for a series. Hence, the result of Lemma 5 is not valid. We would like to thank the community for pointing out this mistake!
This paper has inconsistent results, i.e., we made some failed claims because we did some mistakes for using the test criterion for a series
References in corpus (5)
- ADADELTA: An Adaptive Learning Rate Method
- SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Stochastic Recursive Gradient Algorithm for Nonconvex Optimization
- Theoretical properties of the global optimizer of two layer neural network