1 paper
Long Yang, Zhao Li, Zehong Hu +4
In this paper, we propose a Thompson Sampling algorithm for \emph{unimodal} bandits, where the expected reward is unimodal over the partially ordered arms. To exploit the unimodal…