1 paper
Ben Hambly, Renyuan Xu, Huining Yang
We explore reinforcement learning methods for finding the optimal policy in the linear quadratic regulator (LQR) problem. In particular, we consider the convergence of policy gradi…