1 paper
Simon Buchholz, Jonas M. Kübler, Bernhard Schölkopf
Multi-armed bandits are one of the theoretical pillars of reinforcement learning. Recently, the investigation of quantum algorithms for multi-armed bandit problems was started, and…