3 papers
cs.LG2026
Bridging the Gap Between Average and Discounted TD Learning
Haoxing Tian, Zaiwei Chen, Ioannis Ch. Paschalidis +1
The analysis of Temporal Difference (TD) learning in the average-reward setting faces notable theoretical difficulties because the Bellman operator is not contractive with respect…
cs.LG2024
One-Shot Averaging for Distributed TD() Under Markov Sampling
Haoxing Tian, Ioannis Ch. Paschalidis, Alex Olshevsky
We consider a distributed setup for reinforcement learning, where each agent has a copy of the same Markov Decision Process but transitions are sampled from the corresponding Marko…
cs.LG2023
On the Performance of Temporal Difference Learning With Neural Networks
Haoxing Tian, Ioannis Ch. Paschalidis, Alex Olshevsky
Neural Temporal Difference (TD) Learning is an approximate temporal difference method for policy evaluation that uses a neural network for function approximation. Analysis of Neura…