A Tractable Algorithm For Finite-Horizon Continuous Reinforcement Learning
arXiv:1906.11245 · doi:10.1109/ICoIAS.2019.00018
Abstract
We consider the finite horizon continuous reinforcement learning problem. Our contribution is three-fold. First,we give a tractable algorithm based on optimistic value iteration for the problem. Next,we give a lower bound on regret of order for any algorithm discretizes the state space, improving the previous regret bound of of Ortner and Ryabko \cite{contrl} for the same problem. Next,under the assumption that the rewards and transitions are Hölder Continuous we show that the upper bound on the discretization error is . Finally,we give some simple experiments to validate our propositions.
InProceedings of International Conference on Intelligent Autonomous System, ICOIAS 2019