1 paper
Yuto Tanimoto, Kenji Fukumizu
While many multi-armed bandit algorithms assume that rewards for all arms are constant across rounds, this assumption does not hold in many real-world scenarios. This paper conside…