1 paper
Victor Boone, Bruno Gaujal
In average reward Markov decision processes, state-of-the-art algorithms for regret minimization follow a well-established framework: They are model-based, optimistic and episodic.…