1 paper · 1 filter
Prashant Mehta, Sean Meyn
In recent work it is shown that Q-learning with linear function approximation is stable, in the sense of bounded parameter estimates, under the (ε,I^º)-tamed Gibbs polic…