1 paper
Milos S. Stankovic, Marko Beko, Srdjan S. Stankovic
In this paper we propose several novel distributed gradient-based temporal difference algorithms for multi-agent off-policy learning of linear approximation of the value function i…