3 citations · 3 across the 1 of their papers we have counts for
1 paper
Hugo Penedones, Carlos Riquelme, Damien Vincent +5
We consider the core reinforcement-learning problem of on-policy value function approximation from a batch of trajectory data, and focus on various issues of Temporal Difference (T…