1 paper
Sayak Ray Chowdhury, Aditya Gopalan
We consider online learning for minimizing regret in unknown, episodic Markov decision processes (MDPs) with continuous states and actions. We develop variants of the UCRL and post…