2 papers
stat.ML2026
Minimizing Human Intervention in Online Classification
William Réveillard, Vasileios Saketos, Alexandre Proutiere +1
Training or fine-tuning large language model (LLM)-based systems often requires costly human feedback, yet there is limited understanding of how to minimize such intervention while…
stat.ML2025
Multimodal Bandits: Regret Lower Bounds and Optimal Algorithms
William Réveillard, Richard Combes
We consider a stochastic multi-armed bandit problem with i.i.d. rewards where the expected reward function is multimodal with at most m modes. We propose the first known computatio…