7 papers
Achieving Sample Complexity for Single-Loop Actor-Critic under Minimal Assumptions
Ishaq Hamza, Zaiwei Chen
In this paper, we establish last-iterate convergence rates for off-policy actor--critic methods in reinforcement learning. In particular, under a single-loop, single-timescale impl…
Natural Policy Gradient as Doubly Smoothed Policy Iteration: A Bellman-Operator Framework
Phalguni Nanda, Zaiwei Chen
In this work, we show that natural policy gradient, a core algorithm in reinforcement learning, admits an exact formulation as a smoothed and averaged form of policy iteration. Spe…
Bridging the Gap Between Average and Discounted TD Learning
Haoxing Tian, Zaiwei Chen, Ioannis Ch. Paschalidis +1
The analysis of Temporal Difference (TD) learning in the average-reward setting faces notable theoretical difficulties because the Bellman operator is not contractive with respect…
From Set Convergence to Pointwise Convergence: Finite-Time Guarantees for Average-Reward Q-Learning with Adaptive Stepsizes
Zaiwei Chen, Phalguni Nanda
This work presents the first finite-time analysis for the last-iterate convergence of average-reward -learning with an asynchronous implementation. A key feature of the algorith…
A Minimal-Assumption Analysis of Q-Learning with Time-Varying Policies
Phalguni Nanda, Zaiwei Chen
In this work, we present the first finite-time analysis of Q-learning with time-varying learning policies (i.e., on-policy sampling) for discounted Markov decision processes under…
Natural Hypergradient Descent: Algorithm Design, Convergence Analysis, and Parallel Implementation
Deyi Kong, Zaiwei Chen, Shuzhong Zhang +1
In this work, we propose Natural Hypergradient Descent (NHGD), a new method for solving bilevel optimization problems. To address the computational bottleneck in hypergradient esti…