39 citations · 69 across the 10 of their papers we have counts for
Showing 2020 · cs.LGShow all
3 papers · 2 filters
cs.LG2020
Approximation Benefits of Policy Gradient Methods with Aggregated States
Daniel Russo
Folklore suggests that policy gradient can be more robust to misspecification than its relative, approximate policy iteration. This paper studies the case of state-aggregated repre…
cs.LG2020
On Linear Convergence of Policy Gradient Methods for Finite MDPs
Jalaj Bhandari, Daniel Russo
We revisit the finite time analysis of policy gradient methods in the one of the simplest settings: finite state and action MDPs with a policy class consisting of all stochastic po…
cs.LG2020★ 6 cited
Policy Gradient Optimization of Thompson Sampling Policies
Seungki Min, Ciamac C. Moallemi, Daniel J. Russo
We study the use of policy gradient algorithms to optimize over a class of generalized Thompson sampling policies. Our central insight is to view the posterior parameter sampled by…