activity
20172022
most citedOff-Policy Reinforcement Learning with Delayed Rewards

14 citations · 46 across the 13 of their papers we have counts for

collaborators

35 papers

cs.LG20221 cited

Near-Optimal Regret Bounds for Multi-batch Reinforcement Learning

Zihan Zhang, Yuhang Jiang, Yuan Zhou +1

In this paper, we study the episodic reinforcement learning (RL) problem modeled by finite-horizon Markov Decision Processes (MDPs) with constraint on the number of batches. The mu…

cs.LG2021

Coordinate-wise Control Variates for Deep Policy Gradients

Yuanyi Zhong, Yuan Zhou, Jian Peng

The control variates (CV) method is widely used in policy gradient estimation to reduce the variance of the gradient estimators in practice. A control variate is applied by subtrac…

cs.LG202114 cited

Off-Policy Reinforcement Learning with Delayed Rewards

Beining Han, Zhizhou Ren, Zuofan Wu +2

We study deep reinforcement learning (RL) algorithms with delayed rewards. In many real-world tasks, instant rewards are often not readily accessible or even defined immediately af…

stat.ME2021

TSEC: a framework for online experimentation under experimental constraints

Simon Mak, Yuanshuo Zhou, Lavonne Hoang +1

Thompson sampling is a popular algorithm for solving multi-armed bandit problems, and has been applied in a wide range of applications, from website design to portfolio optimizatio…

cs.LG20203 cited

Harnessing Distribution Ratio Estimators for Learning Agents with Quality and Diversity

Tanmay Gangwani, Jian Peng, Yuan Zhou

Quality-Diversity (QD) is a concept from Neuroevolution with some intriguing applications to Reinforcement Learning. It facilitates learning a population of agents where each membe…

cs.LG2020

Learning Guidance Rewards with Trajectory-space Smoothing

Tanmay Gangwani, Yuan Zhou, Jian Peng

Long-term temporal credit assignment is an important challenge in deep reinforcement learning (RL). It refers to the ability of the agent to attribute actions to consequences that…