1 paper
Yeoneung Kim, Gihun Kim, Jiwhan Park +1
We propose a novel Thompson sampling algorithm that learns linear quadratic regulators (LQR) with a Bayesian regret bound of O(T). Our method leverages Langevin dynamics w…