1 paper
Kaiwen Wang, Owen Oertell, Alekh Agarwal +2
In this paper, we prove that Distributional Reinforcement Learning (DistRL), which learns the return distribution, can obtain second-order bounds in both online and offline RL in g…