4 papers
A Single Stepsize Suffices for Unprojected Linear TD(0): Simultaneous Robust and Fast Rates via Polyak--Ruppert Averaging
Wei-Cheng Lee, Francesco Orabona
We study linear TD(0) under Markovian sampling, where data are generated along a single trajectory. We provide high-probability guarantees for a plain unprojected TD(0) algorithm w…
A Robust Rate for Unprojected TD Learning with Linear Function Approximation
Wei-Cheng Lee, Francesco Orabona
We investigate the finite-time convergence properties of Temporal Difference (TD) learning with linear function approximation, a cornerstone of reinforcement learning. We are inter…
A Best-of-Both-Worlds Proof for Tsallis-INF without Fenchel Conjugates
Wei-Cheng Lee, Francesco Orabona
In this short note, we present a simple derivation of the best-of-both-world guarantee for the Tsallis-INF multi-armed bandit algorithm from J. Zimmert and Y. Seldin. Tsallis-INF:…
New Lower Bounds for Stochastic Non-Convex Optimization through Divergence Decomposition
El Mehdi Saad, Wei-Cheng Lee, Francesco Orabona
We study fundamental limits of first-order stochastic optimization in a range of nonconvex settings, including L-smooth functions satisfying Quasar-Convexity (QC), Quadratic Growth…