1 paper
Hongda Hu, Arthur Charpentier, Mario Ghossoub +1
The classical multi-armed bandit (MAB) problem involves a learner and a collection of K independent arms, each with its own ex ante unknown independent reward distribution. At each…