4 papers
Stability and Sensitivity Analysis of Relative Temporal-Difference Learning: Extended Version
Masoud S. Sakha, Rushikesh Kamalapurkar, Sean Meyn
Relative temporal-difference (TD) learning was introduced to mitigate the slow convergence of TD methods when the discount factor approaches one by subtracting a baseline from the…
Optimistic Training and Convergence of Q-Learning -- Extended Version
Prashant Mehta, Sean Meyn
In recent work it is shown that Q-learning with linear function approximation is stable, in the sense of bounded parameter estimates, under the -tamed Gibbs polic…
Global Convergence and Acceleration for Single Observation Gradient Free Optimization
Caio Kalil Lauand, Sean Meyn
Simultaneous perturbation stochastic approximation (SPSA) is an approach to gradient-free optimization introduced by Spall as a simplification of the approach of Kiefer and Wolfowi…
Revisiting Step-Size Assumptions in Stochastic Approximation
Caio Kalil Lauand, Sean Meyn
Many machine learning and optimization algorithms are built upon the framework of stochastic approximation (SA), for which the selection of step-size (or learning rate) …