1 paper
Jianqing Fan, Zhaoran Wang, Zhuoran Yang +1
We study high-dimensional multi-armed contextual bandits with batched feedback where the T steps of online interactions are divided into L batches. In specific, each batch coll…