6 papers
Last-Iterate Analyses of FTRL with the 1/2-Tsallis Entropy in Stochastic Bandits
Jingxin Zhan, Yuze Han, Zhihua Zhang
The convergence analysis of online learning algorithms is central to machine learning theory, where the last-iterate convergence is particularly important, as it captures the learn…
Accelerated Distributional Temporal Difference Learning with Linear Function Approximation
Kaicheng Jin, Yang Peng, Jiansheng Yang +1
In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The purpose of distributional TD…
Matrix Moment and Concentration Inequalities for Martingales and Ergodic Markov Chains with Applications in Statistical Learning
Yang Peng, Yuchen Xin, Zhihua Zhang
In this paper, we study moment and concentration inequalities for the spectral norm of sums of dependent random matrices. We establish novel Rosenthal-Burkholder inequalities for d…
Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems
Jingxin Zhan, Yuchen Xin, Chenjie Sun +1
We consider a common case of the combinatorial semi-bandit problem, the -set semi-bandit, where the learner exactly selects arms from the total arms. In the adversarial…
A Finite Sample Analysis of Distributional TD Learning with Linear Function Approximation
Yang Peng, Kaicheng Jin, Liangyu Zhang +1
In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The aim of distributional TD lea…
A Regularized Online Newton Method for Stochastic Convex Bandits with Linear Vanishing Noise
Jingxin Zhan, Yuchen Xin, Kaicheng Jin +1
We study a stochastic convex bandit problem where the subgaussian noise parameter is assumed to decrease linearly as the learner selects actions closer and closer to the minimizer…