1 paper
Ruofan Wu, Junmin Zhong, Jennie Si
Policy gradient methods in actor-critic reinforcement learning (RL) have become perhaps the most promising approaches to solving continuous optimal control problems. However, the t…