5 papers
ADDQ: Adaptive Distributional Double Q-Learning
Leif Döring, Benedikt Wille, Maximilian Birr +2
Bias problems in the estimation of -values are a well-known obstacle that slows down convergence of -learning and actor-critic methods. One of the reasons of the success of m…
Almost sure convergence rates of stochastic gradient methods under gradient domination
Simon Weissmann, Sara Klein, Waïss Azizian +1
Stochastic gradient methods are among the most important algorithms in training machine learning problems. While classical assumptions such as strong convexity allow a simple analy…
Clustered KL-barycenter design for policy evaluation
Simon Weissmann, Till Freihaut, Claire Vernade +2
In the context of stochastic bandit models, this article examines how to design sample-efficient behavior policies for the importance sampling evaluation of multiple target policie…
Structure Matters: Dynamic Policy Gradient
Sara Klein, Xiangyuan Zhang, Tamer BaÅar +2
In this work, we study -discounted infinite-horizon tabular Markov decision processes (MDPs) and introduce a framework called dynamic policy gradient (DynPG). The framework dir…
Gradient Span Algorithms Make Predictable Progress in High Dimension
Felix Benning, Leif Döring
We prove that all 'gradient span algorithms' have asymptotically deterministic behavior on scaled Gaussian random functions as the dimension tends to infinity. This is a functional…