1 paper
Tianyuan Jin, Yu Yang, Jing Tang +2
We study the batched best arm identification (BBAI) problem, where the learner's goal is to identify the best arm while switching the policy as less as possible. In particular, we…