33 citations · 33 across the 1 of their papers we have counts for
1 paper
Mohammad Ghavamzadeh, Yaakov Engel, Michal Valko
Policy gradient methods are reinforcement learning algorithms that adapt a parameterized policy by following a performance gradient estimate. Conventional policy gradient methods u…