5 citations · 19 across the 8 of their papers we have counts for
1 paper · 1 filter
Gang Wang, Bingcong Li, Georgios B. Giannakis
Motivated by the widespread use of temporal-difference (TD-) and Q-learning algorithms in reinforcement learning, this paper studies a class of biased stochastic approximation (SA)…