1 paper
Yunwen Guo, Yunlun Shu, Gongyi Zhuo +1
The batched multi-armed bandit (MAB) problem, where rewards are collected in batches, is pivotal in applications like clinical trials. While prior work assumes light-tailed reward…