4 papers · 1 filter
ADDQ: Adaptive Distributional Double Q-Learning
Leif Döring, Benedikt Wille, Maximilian Birr +2
Bias problems in the estimation of -values are a well-known obstacle that slows down convergence of -learning and actor-critic methods. One of the reasons of the success of m…
Almost sure convergence rates of stochastic gradient methods under gradient domination
Simon Weissmann, Sara Klein, Waïss Azizian +1
Stochastic gradient methods are among the most important algorithms in training machine learning problems. While classical assumptions such as strong convexity allow a simple analy…
Clustered KL-barycenter design for policy evaluation
Simon Weissmann, Till Freihaut, Claire Vernade +2
In the context of stochastic bandit models, this article examines how to design sample-efficient behavior policies for the importance sampling evaluation of multiple target policie…
Structure Matters: Dynamic Policy Gradient
Sara Klein, Xiangyuan Zhang, Tamer BaÅar +2
In this work, we study -discounted infinite-horizon tabular Markov decision processes (MDPs) and introduce a framework called dynamic policy gradient (DynPG). The framework dir…