4 papers
Does 1/2-Tsallis-INF Also Work Well for Best-Arm Identification?
Jingxin Zhan, Yuze Han, Zhihua Zhang
Regret minimization (RM) and best-arm identification (BAI) are two fundamental objectives in multi-armed bandits. Among regret-minimizing algorithms, -Tsallis-INF is a canonic…
Last-Iterate Analyses of FTRL with the 1/2-Tsallis Entropy in Stochastic Bandits
Jingxin Zhan, Yuze Han, Zhihua Zhang
The convergence analysis of online learning algorithms is central to machine learning theory, where the last-iterate convergence is particularly important, as it captures the learn…
Follow-the-Perturbed-Leader Approaches Best-of-Both-Worlds for the m-Set Semi-Bandit Problems
Jingxin Zhan, Yuchen Xin, Chenjie Sun +1
We consider a common case of the combinatorial semi-bandit problem, the -set semi-bandit, where the learner exactly selects arms from the total arms. In the adversarial…
A Regularized Online Newton Method for Stochastic Convex Bandits with Linear Vanishing Noise
Jingxin Zhan, Yuchen Xin, Kaicheng Jin +1
We study a stochastic convex bandit problem where the subgaussian noise parameter is assumed to decrease linearly as the learner selects actions closer and closer to the minimizer…