1 paper
Théo Vincent, Kevin Gerhardt, Yogesh Tripathi +5
Temporal-difference (TD) learning is highly effective at controlling and evaluating an agent's long-term outcomes. Most approaches in this paradigm implement a semi-gradient update…