1 paper
Long Yang, Yu Zhang, Gang Zheng +5
Improving sample efficiency has been a longstanding goal in reinforcement learning. This paper proposes VRMPO algorithm: a sample efficient policy gradient method with s…