1 paper
Xiaocan Li, Shiliang Wu, Zheng Shen
Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setting. Decoupled loss used in decoupled P…