7 papers
Efficient Multinomial Logistic Bandit via Frequent Directions
Linzhe He, Yu-Jie Zhang, Sifan Yang +1
This paper studies efficient online algorithms for multinomial logistic bandits (MLogB), where the feedback distribution over outcomes follows a multinomial logistic model of…
Near-Optimal Regret in Adversarial Kernel Bandits
Yu-Jie Zhang, Hao Qiu, Jonathan Scarlett +1
We study the adversarial kernel bandit problem, in which the loss at each round is induced by an arbitrary bounded element of a reproducing kernel Hilbert space (RKHS). We propose…
Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer
Yan-Feng Xie, Yu-Jie Zhang, Peng Zhao +1
We study dynamic regret minimization in non-stationary online learning, with a primary focus on follow-the-regularized-leader (FTRL) methods. FTRL is important for curved losses an…
Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update
Jing Wang, Yu-Jie Zhang, Peng Zhao +1
We study the stochastic linear bandits with heavy-tailed noise. Two principled strategies for handling heavy-tailed noise, truncation and median-of-means, have been introduced to h…
Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation
Long-Fei Li, Yu-Jie Zhang, Peng Zhao +1
We study a new class of MDPs that employs multinomial logit (MNL) function approximation to ensure valid probability distributions over the state space. Despite its significant ben…
Dynamic Regret of Convex and Smooth Functions
Peng Zhao, Yu-Jie Zhang, Lijun Zhang +1
We investigate online convex optimization in non-stationary environments and choose the dynamic regret as the performance measure, defined as the difference between cumulative loss…