2 papers
stat.ML2025
Catoni Contextual Bandits are Robust to Heavy-tailed Rewards
Chenlu Ye, Yujia Jin, Alekh Agarwal +1
Typical contextual bandit algorithms assume that the rewards at each round lie in some fixed range , and their regret scales polynomially with this reward range . Howeve…
stat.ML2023
Provably Efficient High-Dimensional Bandit Learning with Batched Feedbacks
Jianqing Fan, Zhaoran Wang, Zhuoran Yang +1
We study high-dimensional multi-armed contextual bandits with batched feedback where the steps of online interactions are divided into batches. In specific, each batch coll…