14 citations · 14 across the 2 of their papers we have counts for
1 paper · 1 filter
Yi Ouyang, Mukul Gagrani, Ashutosh Nayyar +1
We consider the problem of learning an unknown Markov Decision Process (MDP) that is weakly communicating in the infinite horizon setting. We propose a Thompson Sampling-based rein…