74 citations · 74 across the 1 of their papers we have counts for
1 paper · 1 filter
Jose A. Arjona-Medina, Michael Gillhofer, Michael Widrich +3
We propose RUDDER, a novel reinforcement learning approach for delayed rewards in finite Markov decision processes (MDPs). In MDPs the Q-values are equal to the expected immediate…