4 citations · 4 across the 1 of their papers we have counts for
3 papers
cs.LG2020★ 4 cited
Variance Reduction for Deep Q-Learning using Stochastic Recursive Gradient
Haonan Jia, Xiao Zhang, Jun Xu +4
Deep Q-learning algorithms often suffer from poor gradient estimations with an excessive variance, resulting in unstable training and poor sampling efficiency. Stochastic variance-…
stat.ML2019
Off-policy Learning for Multiple Loggers
Li He, Long Xia, Wei Zeng +3
It is well known that the historical logs are used for evaluating and learning policies in interactive systems, e.g. recommendation, search, and online advertising. Since direct on…
cs.LG2018
MQGrad: Reinforcement Learning of Gradient Quantization in Parameter Server
Guoxin Cui, Jun Xu, Wei Zeng +3
One of the most significant bottleneck in training large scale machine learning models on parameter server (PS) is the communication overhead, because it needs to frequently exchan…