1 paper
Ruiqi Zhang, Yuexiang Zhai, Andrea Zanette
What can an agent learn in a stochastic Multi-Armed Bandit (MAB) problem from a dataset that contains just a single sample for each arm? Surprisingly, in this work, we demonstrate…