1 paper
Akshay Mete, Rahul Singh, Xi Liu +1
The Reward-Biased Maximum Likelihood Estimate (RBMLE) for adaptive control of Markov chains was proposed to overcome the central obstacle of what is variously called the fundamenta…