4 citations · 8 across the 4 of their papers we have counts for
7 papers
Stagewise Enlargement of Batch Size for SGD-based Learning
Shen-Yi Zhao, Yin-Peng Xie, Wu-Jun Li
Existing research shows that the batch size can seriously affect the performance of stochastic gradient descent~(SGD) based learning, including training speed and generalization ab…
Clustered Reinforcement Learning
Xiao Ma, Shen-Yi Zhao, Wu-Jun Li
Exploration strategy design is one of the challenging problems in reinforcement learning~(RL), especially when the environment contains a large state space or sparse rewards. Durin…
ADASS: Adaptive Sample Selection for Training Acceleration
Shen-Yi Zhao, Hao Gao, Wu-Jun Li
Stochastic gradient decent~(SGD) and its variants, including some accelerated variants, have become popular for training in machine learning. However, in all existing SGD and its v…
On the Convergence of Memory-Based Distributed SGD
Shen-Yi Zhao, Hao Gao, Wu-Jun Li
Distributed stochastic gradient descent~(DSGD) has been widely used for optimizing large-scale machine learning models, including both convex and non-convex models. With the rapid…
Quantized Epoch-SGD for Communication-Efficient Distributed Learning
Shen-Yi Zhao, Hao Gao, Wu-Jun Li
Due to its efficiency and ease to implement, stochastic gradient descent (SGD) has been widely used in machine learning. In particular, SGD is one of the most popular optimization…
Proximal SCOPE for Distributed Sparse Learning: Better Data Partition Implies Faster Convergence Rate
Shen-Yi Zhao, Gong-Duo Zhang, Ming-Wei Li +1
Distributed sparse learning with a cluster of multiple machines has attracted much attention in machine learning, especially for large-scale applications with high-dimensional data…