activity
20182022
most citedBeBold: Exploration Beyond the Boundary of Explored Regions

18 citations · 49 across the 7 of their papers we have counts for

collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG20222 cited

Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization

Zihan Zhou, Wei Fu, Bingliang Zhang +1

We present Reward-Switching Policy Optimization (RSPO), a paradigm to discover diverse strategies in complex RL environments by iteratively finding novel policies that are both loc…

cs.LG20212 cited

DAIR: Disentangled Attention Intrinsic Regularization for Safe and Efficient Bimanual Manipulation

Minghao Zhang, Pingcheng Jian, Yi Wu +2

We address the problem of safely solving complex bimanual robot manipulation tasks with sparse rewards. Such challenging tasks can be decomposed into sub-tasks that are accomplisha…

cs.LG20218 cited

Solving Compositional Reinforcement Learning Problems via Task Reduction

Yunfei Li, Yilin Wu, Huazhe Xu +2

We propose a novel learning paradigm, Self-Imitation via Reduction (SIR), for solving compositional reinforcement learning problems. SIR is based on two core ideas: task reduction…

cs.LG202018 cited

BeBold: Exploration Beyond the Boundary of Explored Regions

Tianjun Zhang, Huazhe Xu, Xiaolong Wang +4

Efficient exploration under sparse rewards remains a key challenge in deep reinforcement learning. To guide exploration, previous work makes extensive use of intrinsic reward (IR).…

cs.LG2020

Multi-Agent Collaboration via Reward Attribution Decomposition

Tianjun Zhang, Huazhe Xu, Xiaolong Wang +4

Recent advances in multi-agent reinforcement learning (MARL) have achieved super-human performance in games like Quake 3 and Dota 2. Unfortunately, these techniques require orders-…