activity
20152020
most citedLearning Belief Representations for Imitation Learning in POMDPs

13 citations · 35 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

10 papers · 1 filter

cs.LG20203 cited

Harnessing Distribution Ratio Estimators for Learning Agents with Quality and Diversity

Tanmay Gangwani, Jian Peng, Yuan Zhou

Quality-Diversity (QD) is a concept from Neuroevolution with some intriguing applications to Reinforcement Learning. It facilitates learning a population of agents where each membe…

cs.LG2020

Off-Policy Interval Estimation with Lipschitz Value Iteration

Ziyang Tang, Yihao Feng, Na Zhang +2

Off-policy evaluation provides an essential tool for evaluating the effects of different policies or treatments using only observed data. When applied to high-stakes scenarios such…

cs.LG2020

Learning Guidance Rewards with Trajectory-space Smoothing

Tanmay Gangwani, Yuan Zhou, Jian Peng

Long-term temporal credit assignment is an important challenge in deep reinforcement learning (RL). It refers to the ability of the agent to attribute actions to consequences that…

cs.LG20201 cited

Efficient Competitive Self-Play Policy Optimization

Yuanyi Zhong, Yuan Zhou, Jian Peng

Reinforcement learning from self-play has recently reported many successes. Self-play, where the agents compete with themselves, is often used to generate training data for iterati…

cs.LG20202 cited

Disentangling Controllable Object through Video Prediction Improves Visual Reinforcement Learning

Yuanyi Zhong, Alexander Schwing, Jian Peng

In many vision-based reinforcement learning (RL) problems, the agent controls a movable object in its visual field, e.g., the player's avatar in video games and the robotic arm in…

cs.LG2019

HeteSpaceyWalk: A Heterogeneous Spacey Random Walk for Heterogeneous Information Network Embedding

Yu He, Yangqiu Song, Jianxin Li +3

Heterogeneous information network (HIN) embedding has gained increasing interests recently. However, the current way of random-walk based HIN embedding methods have paid few attent…