10 citations · 42 across the 16 of their papers we have counts for
4 papers · 1 filter
Stagewise Enlargement of Batch Size for SGD-based Learning
Shen-Yi Zhao, Yin-Peng Xie, Wu-Jun Li
Existing research shows that the batch size can seriously affect the performance of stochastic gradient descent~(SGD) based learning, including training speed and generalization ab…
ADASS: Adaptive Sample Selection for Training Acceleration
Shen-Yi Zhao, Hao Gao, Wu-Jun Li
Stochastic gradient decent~(SGD) and its variants, including some accelerated variants, have become popular for training in machine learning. However, in all existing SGD and its v…
On the Convergence of Memory-Based Distributed SGD
Shen-Yi Zhao, Hao Gao, Wu-Jun Li
Distributed stochastic gradient descent~(DSGD) has been widely used for optimizing large-scale machine learning models, including both convex and non-convex models. With the rapid…
Proximal SCOPE for Distributed Sparse Learning: Better Data Partition Implies Faster Convergence Rate
Shen-Yi Zhao, Gong-Duo Zhang, Ming-Wei Li +1
Distributed sparse learning with a cluster of multiple machines has attracted much attention in machine learning, especially for large-scale applications with high-dimensional data…