1 paper
Vasos Arnaoutis, Eric Lutters, Bojana RosiÄ
In this paper, we present a generalized temporal-difference (TD) reinforcement learning framework based on the theory of conditional expectations. The value and action-value (Q-val…