activity
20182020
most citedStagewise Enlargement of Batch Size for SGD-based Learning

4 citations · 8 across the 4 of their papers we have counts for

collaborators

7 papers

stat.ML20204 cited

Stagewise Enlargement of Batch Size for SGD-based Learning

Shen-Yi Zhao, Yin-Peng Xie, Wu-Jun Li

Existing research shows that the batch size can seriously affect the performance of stochastic gradient descent~(SGD) based learning, including training speed and generalization ab…

cs.LG20191 cited

Clustered Reinforcement Learning

Xiao Ma, Shen-Yi Zhao, Wu-Jun Li

Exploration strategy design is one of the challenging problems in reinforcement learning~(RL), especially when the environment contains a large state space or sparse rewards. Durin…

stat.ML2019

ADASS: Adaptive Sample Selection for Training Acceleration

Shen-Yi Zhao, Hao Gao, Wu-Jun Li

Stochastic gradient decent~(SGD) and its variants, including some accelerated variants, have become popular for training in machine learning. However, in all existing SGD and its v…

stat.ML20191 cited

On the Convergence of Memory-Based Distributed SGD

Shen-Yi Zhao, Hao Gao, Wu-Jun Li

Distributed stochastic gradient descent~(DSGD) has been widely used for optimizing large-scale machine learning models, including both convex and non-convex models. With the rapid…

cs.LG20192 cited

Quantized Epoch-SGD for Communication-Efficient Distributed Learning

Shen-Yi Zhao, Hao Gao, Wu-Jun Li

Due to its efficiency and ease to implement, stochastic gradient descent (SGD) has been widely used in machine learning. In particular, SGD is one of the most popular optimization…

stat.ML2018

Proximal SCOPE for Distributed Sparse Learning: Better Data Partition Implies Faster Convergence Rate

Shen-Yi Zhao, Gong-Duo Zhang, Ming-Wei Li +1

Distributed sparse learning with a cluster of multiple machines has attracted much attention in machine learning, especially for large-scale applications with high-dimensional data…