1 paper
Yukinari Hisaki, Isao Ono
In this paper, we propose an off-policy deep reinforcement learning (DRL) method utilizing the average reward criterion. While most existing DRL methods employ the discounted rewar…