1 paper
Joery A. de Vries, Jinke He, Yaniv Oren +3
Optimally trading-off exploration and exploitation is the holy grail of reinforcement learning as it promises maximal data-efficiency for solving any task. Bayes-optimal agents ach…