1 paper
Wesley A. Suttle, Amrit Singh Bedi, Bhrij Patel +3
Many existing reinforcement learning (RL) methods employ stochastic gradient iteration on the back end, whose stability hinges upon a hypothesis that the data-generating process mi…