1 paper
Jonatha Anselmi, Bruno Gaujal, Louis-Sébastien Rebuffi
In this paper, we revisit the regret of undiscounted reinforcement learning in MDPs with a birth and death structure. Specifically, we consider a controlled queue with impatient jo…