2 citations · 2 across the 1 of their papers we have counts for
1 paper
Chengqian Gao, Ke Xu, Liu Liu +3
A promising paradigm for offline reinforcement learning (RL) is to constrain the learned policy to stay close to the dataset behaviors, known as policy constraint offline RL. Howev…