Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Asymptotically Optimal Problem-Dependent Bandit Policies for Transfer Learning
Adrien Prevost, Timothee Mathieu, Odalric-Ambrym Maillard
We study the non-contextual multi-armed bandit problem in a transfer learning setting: before any pulls, the learner is given N'_k i.i.d. samples from each source distribution nu'_…
cs.LG2024
AdaStop: adaptive statistical testing for sound comparisons of Deep RL agents
Timothée Mathieu, Riccardo Della Vecchia, Alena Shilova +4
Recently, the scientific community has questioned the statistical reproducibility of many empirical results, especially in the field of machine learning. To contribute to the resol…