1 citations · 1 across the 1 of their papers we have counts for
1 paper
Kai-Chun Hu, Chen-Huan Pi, Ting Han Wei +4
In this paper, we point out a fundamental property of the objective in reinforcement learning, with which we can reformulate the policy gradient objective into a perceptron-like lo…