activity
20182021
most citedDistributional Reinforcement Learning for Efficient Exploration

29 citations · 30 across the 2 of their papers we have counts for

collaborators

11 papers

cs.LG20211 cited

Learning Expected Emphatic Traces for Deep RL

Ray Jiang, Shangtong Zhang, Veronica Chelu +2

Off-policy sampling and experience replay are key for improving sample efficiency and scaling model-free temporal difference learning methods. When combined with function approxima…

cs.LG2020

Learning Retrospective Knowledge with Reverse Reinforcement Learning

Shangtong Zhang, Vivek Veeriah, Shimon Whiteson

We present a Reverse Reinforcement Learning (Reverse RL) approach for representing retrospective knowledge. General Value Functions (GVFs) have enjoyed great success in representin…

cs.LG2020

GradientDICE: Rethinking Generalized Offline Estimation of Stationary Values

Shangtong Zhang, Bo Liu, Shimon Whiteson

We present GradientDICE for estimating the density ratio between the state distribution of the target policy and the sampling distribution in off-policy reinforcement learning. Gra…

cs.LG2019

Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation

Shangtong Zhang, Bo Liu, Hengshuai Yao +1

We present the first provably convergent two-timescale off-policy actor-critic algorithm (COF-PAC) with function approximation. Key to COF-PAC is the introduction of a new critic,…

cs.LG201929 cited

Distributional Reinforcement Learning for Efficient Exploration

Borislav Mavrin, Shangtong Zhang, Hengshuai Yao +3

In distributional reinforcement learning (RL), the estimated distribution of value function models both the parametric and intrinsic uncertainties. We propose a novel and efficient…

cs.AI2019

Mega-Reward: Achieving Human-Level Play without Extrinsic Rewards

Yuhang Song, Jianyi Wang, Thomas Lukasiewicz +4

Intrinsic rewards were introduced to simulate how human intelligence works; they are usually evaluated by intrinsically-motivated play, i.e., playing games without extrinsic reward…