3 papers
cs.LG2021
Online Sign Identification: Minimization of the Number of Errors in Thresholding Bandits
Reda Ouhamma, Rémy Degenne, Pierre Gaillard +1
In the fixed budget thresholding bandit problem, an algorithm sequentially allocates a budgeted number of samples to different distributions. It then predicts whether the mean of e…
cs.LG2018
Bridging the gap between regret minimization and best arm identification, with application to A/B tests
Rémy Degenne, Thomas Nedelec, Clément Calauzènes +1
State of the art online learning procedures focus either on selecting the best alternative ("best arm identification") or on minimizing the cost (the "regret"). We merge these two…
cs.LG2018
Bandits with Side Observations: Bounded vs. Logarithmic Regret
Rémy Degenne, Evrard Garcelon, Vianney Perchet
We consider the classical stochastic multi-armed bandit but where, from time to time and roughly with frequency , an extra observation is gathered by the agent for free. We prov…