1 paper
Lorenzo Croissant, Marc Abeille, Bruno Bouchard
We consider the Reinforcement Learning problem of controlling an unknown dynamical system to maximise the long-term average reward along a single trajectory. Most of the literature…