1 paper · 1 filter
Philipp Altmann, Thomy Phan, Fabian Ritz +2
We propose discriminative reward co-training (DIRECT) as an extension to deep reinforcement learning algorithms. Building upon the concept of self-imitation learning (SIL), we intr…