1 paper
Sanjeev Manivannan, Shuban V
We address the discounted reward setting in reinforcement learning (RL). To mitigate the value approximation challenges in policy gradient methods, actor-critic approaches have bee…