3 papers
cs.LG2026
Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework
Phalguni Nanda, Zaiwei Chen
In this work, we show that natural policy gradient, a core algorithm in reinforcement learning, admits an exact formulation as a smoothed and averaged form of policy iteration. Spe…
cs.LG2026
From Set Convergence to Pointwise Convergence: Finite-Time Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes
Zaiwei Chen, Phalguni Nanda
This work presents the first finite-time analysis for the last-iterate convergence of average-reward -learning with an asynchronous implementation. A key feature of the algorith…
cs.LG2026
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
Phalguni Nanda, Zaiwei Chen
In this work, we present the first finite-time analysis of Q-learning with time-varying learning policies (i.e., on-policy sampling) for discounted Markov decision processes under…