3 papers
cs.LG2026
Leveraging Similarities in Multi-Armed Bandits
Khaled Eldowa, Thibaud Rahier, Augustin Cablant +2
In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags, or hierarchical structure.…
stat.ML2026
Functional Natural Policy Gradients
Aurelien Bibaut, Houssam Zenati, Thibaud Rahier +1
We propose a cross-fitted debiasing device for policy learning from offline data. A key consequence of the resulting learning principle is regret even for policy classes…
cs.LG2025
Logarithmic Regret for Unconstrained Submodular Maximization Stochastic Bandit
Julien Zhou, Pierre Gaillard, Thibaud Rahier +1
We address the online unconstrained submodular maximization problem (Online USM), in a setting with stochastic bandit feedback. In this framework, a decision-maker receives noisy r…