29 citations · 30 across the 2 of their papers we have counts for
11 papers
Learning Expected Emphatic Traces for Deep RL
Ray Jiang, Shangtong Zhang, Veronica Chelu +2
Off-policy sampling and experience replay are key for improving sample efficiency and scaling model-free temporal difference learning methods. When combined with function approxima…
Learning Retrospective Knowledge with Reverse Reinforcement Learning
Shangtong Zhang, Vivek Veeriah, Shimon Whiteson
We present a Reverse Reinforcement Learning (Reverse RL) approach for representing retrospective knowledge. General Value Functions (GVFs) have enjoyed great success in representin…
GradientDICE: Rethinking Generalized Offline Estimation of Stationary Values
Shangtong Zhang, Bo Liu, Shimon Whiteson
We present GradientDICE for estimating the density ratio between the state distribution of the target policy and the sampling distribution in off-policy reinforcement learning. Gra…
Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation
Shangtong Zhang, Bo Liu, Hengshuai Yao +1
We present the first provably convergent two-timescale off-policy actor-critic algorithm (COF-PAC) with function approximation. Key to COF-PAC is the introduction of a new critic,…
Distributional Reinforcement Learning for Efficient Exploration
Borislav Mavrin, Shangtong Zhang, Hengshuai Yao +3
In distributional reinforcement learning (RL), the estimated distribution of value function models both the parametric and intrinsic uncertainties. We propose a novel and efficient…
Mega-Reward: Achieving Human-Level Play without Extrinsic Rewards
Yuhang Song, Jianyi Wang, Thomas Lukasiewicz +4
Intrinsic rewards were introduced to simulate how human intelligence works; they are usually evaluated by intrinsically-motivated play, i.e., playing games without extrinsic reward…