42 citations · 56 across the 3 of their papers we have counts for
3 papers
cs.LG2021
Learning Domain Invariant Representations in Goal-conditioned Block MDPs
Beining Han, Chongyi Zheng, Harris Chan +3
Deep Reinforcement Learning (RL) is successful in solving many complex Markov Decision Processes (MDPs) problems. However, agents often face unanticipated environmental changes aft…
cs.LG2021★ 14 cited
Off-Policy Reinforcement Learning with Delayed Rewards
Beining Han, Zhizhou Ren, Zuofan Wu +2
We study deep reinforcement learning (RL) algorithms with delayed rewards. In many real-world tasks, instant rewards are often not readily accessible or even defined immediately af…
cs.LG2020★ 42 cited
Off-Policy Multi-Agent Decomposed Policy Gradients
Yihan Wang, Beining Han, Tonghan Wang +2
Multi-agent policy gradient (MAPG) methods recently witness vigorous progress. However, there is a significant performance discrepancy between MAPG methods and state-of-the-art mul…