18 citations · 49 across the 7 of their papers we have counts for
5 papers · 1 filter
Continuously Discovering Novel Strategies via Reward-Switching Policy Optimization
Zihan Zhou, Wei Fu, Bingliang Zhang +1
We present Reward-Switching Policy Optimization (RSPO), a paradigm to discover diverse strategies in complex RL environments by iteratively finding novel policies that are both loc…
DAIR: Disentangled Attention Intrinsic Regularization for Safe and Efficient Bimanual Manipulation
Minghao Zhang, Pingcheng Jian, Yi Wu +2
We address the problem of safely solving complex bimanual robot manipulation tasks with sparse rewards. Such challenging tasks can be decomposed into sub-tasks that are accomplisha…
Solving Compositional Reinforcement Learning Problems via Task Reduction
Yunfei Li, Yilin Wu, Huazhe Xu +2
We propose a novel learning paradigm, Self-Imitation via Reduction (SIR), for solving compositional reinforcement learning problems. SIR is based on two core ideas: task reduction…
BeBold: Exploration Beyond the Boundary of Explored Regions
Tianjun Zhang, Huazhe Xu, Xiaolong Wang +4
Efficient exploration under sparse rewards remains a key challenge in deep reinforcement learning. To guide exploration, previous work makes extensive use of intrinsic reward (IR).…
Multi-Agent Collaboration via Reward Attribution Decomposition
Tianjun Zhang, Huazhe Xu, Xiaolong Wang +4
Recent advances in multi-agent reinforcement learning (MARL) have achieved super-human performance in games like Quake 3 and Dota 2. Unfortunately, these techniques require orders-…