1 paper · 1 filter
Ahmad Ahmad, Mehdi Kermanshah, Kevin Leahy +6
In this paper, we tackle the challenging problem of delayed rewards in reinforcement learning (RL). While Proximal Policy Optimization (PPO) has emerged as a leading Policy Gradien…