3 papers
stat.ML2026
Minimizing Human Intervention in Online Classification
William Réveillard, Vasileios Saketos, Alexandre Proutiere +1
Training or fine-tuning large language model (LLM)-based systems often requires costly human feedback, yet there is limited understanding of how to minimize such intervention while…
stat.ML2025
Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms
William Réveillard, Richard Combes
We consider a stochastic multi-armed bandit problem with i.i.d. rewards where the expected reward function is multimodal with at most m modes. We propose the first known computatio…
cs.LG2024
Low-Rank Bandits via Tight Two-to-Infinity Singular Subspace Recovery
Yassir Jedra, William Réveillard, Stefan Stojanovic +1
We study contextual bandits with low-rank structure where, in each round, if the (context, arm) pair is selected, the learner observes a noisy sample of th…