11 citations · 11 across the 1 of their papers we have counts for
1 paper
Jialong Wu, Haixu Wu, Zihan Qiu +2
Policy constraint methods to offline reinforcement learning (RL) typically utilize parameterization or regularization that constrains the policy to perform actions within the suppo…