1 paper
Qing Wang, Yingru Li, Jiechao Xiong +1
In deep reinforcement learning, policy optimization methods need to deal with issues such as function approximation and the reuse of off-policy data. Standard policy gradient metho…