3 papers
cs.LG2026
The Role of Target Update Frequencies in Q-Learning
Simon Weissmann, Tilman Aach, Benedikt Wille +2
The target network update frequency (TUF) is a central stabilization mechanism in (deep) Q-learning. However, their selection remains poorly understood and is often treated merely…
cs.LG2026
An Approximate Ascent Approach To Prove Convergence of PPO
Leif Doering, Daniel Schmidt, Moritz Melcher +4
Proximal Policy Optimization (PPO) is among the most widely used deep reinforcement learning algorithms, yet its theoretical foundations remain incomplete. Most importantly, conver…
cs.LG2025
ADDQ: Adaptive Distributional Double Q-Learning
Leif Döring, Benedikt Wille, Maximilian Birr +2
Bias problems in the estimation of -values are a well-known obstacle that slows down convergence of -learning and actor-critic methods. One of the reasons of the success of m…