1 paper
Lucas Weber, Ana BuÅ¡iÄ, Jiamin Zhu
The expected regret of any reinforcement learning algorithm is lower bounded by I^c◯(DXAT) for undiscounted returns, where D is the diameter of the Markov decis…