1 paper
Abhishek Naik, Yi Wan, Manan Tomar +1
We show that discounted methods for solving continuing reinforcement learning problems can perform significantly better if they center their rewards by subtracting out the rewards'…