14 citations · 46 across the 13 of their papers we have counts for
35 papers
Near-Optimal Regret Bounds for Multi-batch Reinforcement Learning
Zihan Zhang, Yuhang Jiang, Yuan Zhou +1
In this paper, we study the episodic reinforcement learning (RL) problem modeled by finite-horizon Markov Decision Processes (MDPs) with constraint on the number of batches. The mu…
Coordinate-wise Control Variates for Deep Policy Gradients
Yuanyi Zhong, Yuan Zhou, Jian Peng
The control variates (CV) method is widely used in policy gradient estimation to reduce the variance of the gradient estimators in practice. A control variate is applied by subtrac…
Off-Policy Reinforcement Learning with Delayed Rewards
Beining Han, Zhizhou Ren, Zuofan Wu +2
We study deep reinforcement learning (RL) algorithms with delayed rewards. In many real-world tasks, instant rewards are often not readily accessible or even defined immediately af…
TSEC: a framework for online experimentation under experimental constraints
Simon Mak, Yuanshuo Zhou, Lavonne Hoang +1
Thompson sampling is a popular algorithm for solving multi-armed bandit problems, and has been applied in a wide range of applications, from website design to portfolio optimizatio…
Harnessing Distribution Ratio Estimators for Learning Agents with Quality and Diversity
Tanmay Gangwani, Jian Peng, Yuan Zhou
Quality-Diversity (QD) is a concept from Neuroevolution with some intriguing applications to Reinforcement Learning. It facilitates learning a population of agents where each membe…
Learning Guidance Rewards with Trajectory-space Smoothing
Tanmay Gangwani, Yuan Zhou, Jian Peng
Long-term temporal credit assignment is an important challenge in deep reinforcement learning (RL). It refers to the ability of the agent to attribute actions to consequences that…