Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
A Single Stepsize Suffices for Unprojected Linear TD(0): Simultaneous Robust and Fast Rates via Polyak--Ruppert Averaging
Wei-Cheng Lee, Francesco Orabona
We study linear TD(0) under Markovian sampling, where data are generated along a single trajectory. We provide high-probability guarantees for a plain unprojected TD(0) algorithm w…
cs.LG2025
A Best-of-Both-Worlds Proof for Tsallis-INF without Fenchel Conjugates
Wei-Cheng Lee, Francesco Orabona
In this short note, we present a simple derivation of the best-of-both-world guarantee for the Tsallis-INF multi-armed bandit algorithm from J. Zimmert and Y. Seldin. Tsallis-INF:…
cs.LG2025
A Robust Rate for Unprojected TD Learning with Linear Function Approximation
Wei-Cheng Lee, Francesco Orabona
We investigate the finite-time convergence properties of Temporal Difference (TD) learning with linear function approximation, a cornerstone of reinforcement learning. We are inter…