1 paper
Miquel Junyent, Anders Jonsson, Vicenç Gómez
Optimal action selection in decision problems characterized by sparse, delayed rewards is still an open challenge. For these problems, current deep reinforcement learning methods r…