1 paper · 1 filter
Diogo S. Carvalho, Pedro A. Santos, Francisco S. Melo
We study the convergence of Q-learning with linear function approximation. Our key contribution is the introduction of a novel multi-Bellman operator that extends the traditional…