activity
20232026
collaborators

10 papers

cs.LG2026

Does 1/2-Tsallis-INF Also Work Well for Best-Arm Identification?

Jingxin Zhan, Yuze Han, Zhihua Zhang

Regret minimization (RM) and best-arm identification (BAI) are two fundamental objectives in multi-armed bandits. Among regret-minimizing algorithms, -Tsallis-INF is a canonic…

cs.LG2026

Active-Trace Complexity Bounds for Moreau--Yosida Unadjusted Langevin Sampling

Yuchen Xin, Zhihua Zhang

We study the Moreau--Yosida unadjusted Langevin algorithm (MYULA) for the nonsmooth composite target \[ π(dx)\propto \exp\{-f(x)-g(x)\}\,dx, \qquad x\in\mathbb R^d, \] where \(f\)…

stat.ML2025

Accelerated Distributional Temporal Difference Learning with Linear Function Approximation

Kaicheng Jin, Yang Peng, Jiansheng Yang +1

In this paper, we study the finite-sample statistical rates of distributional temporal difference (TD) learning with linear function approximation. The purpose of distributional TD…

cs.LG2025

Last-Iterate Analyses of FTRL with the 1/2-Tsallis Entropy in Stochastic Bandits

Jingxin Zhan, Yuze Han, Zhihua Zhang

The convergence analysis of online learning algorithms is central to machine learning theory, where the last-iterate convergence is particularly important, as it captures the learn…

math.PR2025

Matrix Moment and Concentration Inequalities for Martingales and Ergodic Markov Chains with Applications in Statistical Learning

Yang Peng, Yuchen Xin, Zhihua Zhang

In this paper, we study moment and concentration inequalities for the spectral norm of sums of dependent random matrices. We establish novel Rosenthal-Burkholder inequalities for t…

cs.LG2025

Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems

Jingxin Zhan, Yuchen Xin, Chenjie Sun +1

We consider a common case of the combinatorial semi-bandit problem, the -set semi-bandit, where the learner exactly selects arms from the total arms. In the adversarial…