1 paper
Simon Schmitt, John Shawe-Taylor, Hado van Hasselt
How to efficiently explore in reinforcement learning is an open problem. Many exploration algorithms employ the epistemic uncertainty of their own value predictions -- for instance…