1 paper
Arun Verma, Zhongxiang Dai, Yao Shu +1
We study a novel variant of the parameterized bandits problem in which the learner can observe additional auxiliary feedback that is correlated with the observed reward. The auxili…