:An Unbiased Stratified Statistic and a Fast Gradient Optimization Algorithm Based on It
arXiv:2110.03354
Abstract
-The fluctuation effect of gradient expectation and variance caused by parameter update between consecutive iterations is neglected or confusing by current mainstream gradient optimization algorithms. The work in this paper remedy this issue by introducing a novel unbiased stratified statistic \ \ , a sufficient condition of fast convergence for \ \ also is established. A novel algorithm named MSSG designed based on \ \ outperforms other sgd-like algorithms. Theoretical conclusions and experimental evidence strongly suggest to employ MSSG when training deep model.
References in corpus (6)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- ADADELTA: An Adaptive Learning Rate Method
- On the Convergence of Adam and Beyond
- SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives
- A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
- A Novel Stochastic Stratified Average Gradient Method: Convergence Rate and Its Complexity