3 papers
cs.LG2026
Optimistic Training and Convergence of Q-Learning -- Extended Version
Prashant Mehta, Sean Meyn
In recent work it is shown that Q-learning with linear function approximation is stable, in the sense of bounded parameter estimates, under the -tamed Gibbs policy…
math.OC2025
Global Convergence and Acceleration for Single Observation Gradient Free Optimization
Caio Kalil Lauand, Sean Meyn
Simultaneous perturbation stochastic approximation (SPSA) is an approach to gradient-free optimization introduced by Spall as a simplification of the approach of Kiefer and Wolfowi…
math.OC2024
Markovian Foundations for Quasi-Stochastic Approximation in Two Timescales: Extended Version
Caio Kalil Lauand, Sean Meyn
Many machine learning and optimization algorithms can be cast as instances of stochastic approximation (SA). The convergence rate of these algorithms is known to be slow, with the…