1 paper
Alan Chan, Kris de Asis, Richard S. Sutton
Value-based methods for reinforcement learning lack generally applicable ways to derive behavior from a value function. Many approaches involve approximate value iteration (e.g., $…