3 papers
cs.LG2019
WALL-E: An Efficient Reinforcement Learning Research Framework
Tianbing Xu, Andrew Zhang, Liang Zhao
There are two halves to RL systems: experience collection time and policy learning time. For a large number of samples in rollouts, experience collection time is the major bottlene…
cs.LG2018
Learning to Explore with Meta-Policy Gradient
Tianbing Xu, Qiang Liu, Liang Zhao +1
The performance of off-policy learning, including deep Q-learning and deep deterministic policy gradient (DDPG), critically depends on the choice of the exploration policy. Existin…
cs.LG2018
Variational Inference for Policy Gradient
Tianbing Xu
Inspired by the seminal work on Stein Variational Inference and Stein Variational Policy Gradient, we derived a method to generate samples from the posterior variational parameter…