1 paper
Abdul Wahab, Raksha Kumaraswamy, Martha White
Optimistic value estimates provide one mechanism for directed exploration in reinforcement learning (RL). The agent acts greedily with respect to an estimate of the value plus what…