3 papers
math.OC2026
Beyond the Bellman Fixed Point: Geometry and Fast Policy Identification in Value Iteration
Donghwan Lee
Q-value iteration (Q-VI) is usually analyzed through the \(γ\)-contraction of the Bellman operator. This argument proves convergence to \(Q^*\), but it gives only a coarse account…
cs.LG2026
Contraction-Aligned Analysis of Soft Bellman Residual Minimization with Weighted Lp-Norm for Markov Decision Problem
Hyukjun Yang, Han-Dong Lim, Donghwan Lee
The problem of solving Markov decision processes under function approximation remains a fundamental challenge, even under linear function approximation settings. A key difficulty a…
cs.LG2026
Learning the Model While Learning Q: Finite-Time Sample Complexity of Online SyncMBQ
Han-Dong Lim, HyeAnn Lee, Donghwan Lee
Reinforcement learning has witnessed significant advancements, particularly with the emergence of model-based approaches. Among these, -learning has proven to be a powerful algo…