Distributed Reinforcement Learning via Gossip
arXiv:1310.7610
Abstract
We consider the classical TD(0) algorithm implemented on a network of agents wherein the agents also incorporate the updates received from neighboring agents using a gossip-like mechanism. The combined scheme is shown to converge for both discounted and average cost problems.
18 pages, 3 figures, Submitted to Discrete Event Dynamic Systems