1 paper
Wensong Bai, Chao Zhang, Yichao Fu +3
In this paper, we propose the first fully push-forward-based distributional reinforcement learning algorithm, named PACER, which consists of a distributional critic, a stochastic a…