1 paper
Xubo Lyu, Site Li, Seth Siriya +2
In this paper, a novel optimal control-based baseline function is presented for the policy gradient method in deep reinforcement learning (RL). The baseline is obtained by computin…