1 paper · 1 filter
Ofir Schlisselberg, Ido Cohen, Tal Lancewicki +1
In this paper, we investigate a variant of the classical stochastic Multi-armed Bandit (MAB) problem, where the payoff received by an agent (either cost or reward) is both delayed,…