1 paper
Long Chen, Yinkui Liu, Shen Li +2
Pseudo-count is an effective anti-exploration method in offline reinforcement learning (RL) by counting state-action pairs and imposing a large penalty on rare or unseen state-acti…