49 citations · 83 across the 7 of their papers we have counts for
1 paper · 1 filter
Nan Ding, Radu Soricut
Policy-gradient approaches to reinforcement learning have two common and undesirable overhead procedures, namely warm-start training and sample variance reduction. In this paper, w…