1 paper
Mark Towers, Yali Du, Christopher Freeman +1
Future reward estimation is a core component of reinforcement learning agents; i.e., Q-value and state-value functions, predicting an agent's sum of future rewards. Their scalar ou…