10 citations · 43 across the 18 of their papers we have counts for
Showing 2020 · stat.MLShow all
2 papers · 2 filters
stat.ML2020
Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training
Shen-Yi Zhao, Chang-Wei Shi, Yin-Peng Xie +1
Stochastic gradient descent~(SGD) and its variants have been the dominating optimization methods in machine learning. Compared to SGD with small-batch training, SGD with large-batc…
stat.ML2020★ 4 cited
Stagewise Enlargement of Batch Size for SGD-based Learning
Shen-Yi Zhao, Yin-Peng Xie, Wu-Jun Li
Existing research shows that the batch size can seriously affect the performance of stochastic gradient descent~(SGD) based learning, including training speed and generalization ab…