1 paper
Abdelkrim Zitouni, Mehdi Hennequin, Juba Agoun +3
We derive a novel PAC-Bayesian generalization bound for reinforcement learning that explicitly accounts for Markov dependencies in the data, through the chain's mixing time. This c…