1 paper
Danil Provodin, Pratik Gajane, Mykola Pechenizkiy +1
We present a new algorithm based on posterior sampling for learning in constrained Markov decision processes (CMDP) in the infinite-horizon undiscounted setting. The algorithm achi…