354 citations · 740 across the 11 of their papers we have counts for
1 paper · 1 filter
Dilip Arumugam, Satinder Singh
The Bayes-Adaptive Markov Decision Process (BAMDP) formalism pursues the Bayes-optimal solution to the exploration-exploitation trade-off in reinforcement learning. As the computat…