7 citations · 7 across the 1 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2021★ 7 cited
Off-Policy Imitation Learning from Observations
Zhuangdi Zhu, Kaixiang Lin, Bo Dai +1
Learning from Observations (LfO) is a practical reinforcement learning scenario from which many applications can benefit through the reuse of incomplete resources. Compared to conv…
cs.LG2020
Learning Sparse Rewarded Tasks from Sub-Optimal Demonstrations
Zhuangdi Zhu, Kaixiang Lin, Bo Dai +1
Model-free deep reinforcement learning (RL) has demonstrated its superiority on many complex sequential decision-making problems. However, heavy dependence on dense rewards and hig…