1 paper
Nithyanand Kota, Abhishek Mishra, Sunil Srinivasa +3
The high variance issue in unbiased policy-gradient methods such as VPG and REINFORCE is typically mitigated by adding a baseline. However, the baseline fitting itself suffers from…