1 paper
Chen-Xiao Gao, Chenyang Wu, Mingjun Cao +3
Behavior regularization, which constrains the policy to stay close to some behavior policy, is widely used in offline reinforcement learning (RL) to manage the risk of hazardous ex…