activity
20192026
collaborators

7 papers

cs.LG2026

Efficient Multinomial Logistic Bandit via Frequent Directions

Linzhe He, Yu-Jie Zhang, Sifan Yang +1

This paper studies efficient online algorithms for multinomial logistic bandits (MLogB), where the feedback distribution over outcomes follows a multinomial logistic model of…

cs.LG2026

Near-Optimal Regret in Adversarial Kernel Bandits

Yu-Jie Zhang, Hao Qiu, Jonathan Scarlett +1

We study the adversarial kernel bandit problem, in which the loss at each round is induced by an arbitrary bounded element of a reproducing kernel Hilbert space (RKHS). We propose…

cs.LG2026

Dynamic Regret via Discounted-to-Dynamic Reduction with Applications to Curved Losses and Adam Optimizer

Yan-Feng Xie, Yu-Jie Zhang, Peng Zhao +1

We study dynamic regret minimization in non-stationary online learning, with a primary focus on follow-the-regularized-leader (FTRL) methods. FTRL is important for curved losses an…

cs.LG2025

Heavy-Tailed Linear Bandits: Huber Regression with One-Pass Update

Jing Wang, Yu-Jie Zhang, Peng Zhao +1

We study the stochastic linear bandits with heavy-tailed noise. Two principled strategies for handling heavy-tailed noise, truncation and median-of-means, have been introduced to h…

cs.LG2024

Provably Efficient Reinforcement Learning with Multinomial Logit Function Approximation

Long-Fei Li, Yu-Jie Zhang, Peng Zhao +1

We study a new class of MDPs that employs multinomial logit (MNL) function approximation to ensure valid probability distributions over the state space. Despite its significant ben…

cs.LG2020

Dynamic Regret of Convex and Smooth Functions

Peng Zhao, Yu-Jie Zhang, Lijun Zhang +1

We investigate online convex optimization in non-stationary environments and choose the dynamic regret as the performance measure, defined as the difference between cumulative loss…