5 papers · 1 filter
The Role of Target Update Frequencies in Q-Learning
Simon Weissmann, Tilman Aach, Benedikt Wille +2
The target network update frequency (TUF) is a central stabilization mechanism in (deep) Q-learning. However, their selection remains poorly understood and is often treated merely…
An Approximate Ascent Approach To Prove Convergence of PPO
Leif Doering, Daniel Schmidt, Moritz Melcher +4
Proximal Policy Optimization (PPO) is among the most widely used deep reinforcement learning algorithms, yet its theoretical foundations remain incomplete. Most importantly, conver…
Almost sure convergence rates of stochastic gradient methods under gradient domination
Simon Weissmann, Sara Klein, Waïss Azizian +1
Stochastic gradient methods are among the most important algorithms in training machine learning problems. While classical assumptions such as strong convexity allow a simple analy…
Clustered KL-barycenter design for policy evaluation
Simon Weissmann, Till Freihaut, Claire Vernade +2
In the context of stochastic bandit models, this article examines how to design sample-efficient behavior policies for the importance sampling evaluation of multiple target policie…
Structure Matters: Dynamic Policy Gradient
Sara Klein, Xiangyuan Zhang, Tamer BaÅar +2
In this work, we study -discounted infinite-horizon tabular Markov decision processes (MDPs) and introduce a framework called dynamic policy gradient (DynPG). The framework dir…