11 citations · 23 across the 5 of their papers we have counts for
1 paper · 1 filter
Matthia Sabatelli, Gilles Louppe, Pierre Geurts +1
We introduce a novel Deep Reinforcement Learning (DRL) algorithm called Deep Quality-Value (DQV) Learning. DQV uses temporal-difference learning to train a Value neural network and…