160 citations · 160 across the 1 of their papers we have counts for
1 paper
Lex Weaver, Nigel Tao
There exist a number of reinforcement learning algorithms which learnby climbing the gradient of expected reward. Their long-runconvergence has been proved, even in partially obser…