1 paper
Chaiwon Kim, Jongyeong Lee, Min-hwan Oh
We study the decoupled multi-armed bandit problem, where the learner separately selects one arm for exploration and one, possibly different, arm for exploitation at each round. In…