1 citations · 1 across the 14 of their papers we have counts for
1 paper · 2 filters
Siqiao Mu, Diego Klabjan
Since the objective functions of reinforcement learning problems are typically highly nonconvex, it is desirable that policy gradient, the most popular algorithm, escapes saddle po…