1 paper
Zhenyu Zhang, Xiangfeng Luo, Tong Liu +5
Instability and slowness are two main problems in deep reinforcement learning. Even if proximal policy optimization (PPO) is the state of the art, it still suffers from these two p…