1 paper
Harshit Dhankhar, Kshitij Mishra, Tejas Bodas
In the realm of multi-arm bandit problems, the Gittins index policy is known to be optimal in maximizing the expected total discounted reward obtained from pulling the Markovian ar…