3 papers
stat.ML2025
Counterfactual Learning of Stochastic Policies with Continuous Actions
Houssam Zenati, Alberto Bietti, Matthieu Martin +3
Counterfactual reasoning from logged data has become increasingly important for many applications such as web advertising or healthcare. In this paper, we address the problem of le…
cs.LG2025
Logarithmic Regret for Unconstrained Submodular Maximization Stochastic Bandit
Julien Zhou, Pierre Gaillard, Thibaud Rahier +1
We address the online unconstrained submodular maximization problem (Online USM), in a setting with stochastic bandit feedback. In this framework, a decision-maker receives noisy r…
cs.LG2024
Towards Efficient and Optimal Covariance-Adaptive Algorithms for Combinatorial Semi-Bandits
Julien Zhou, Pierre Gaillard, Thibaud Rahier +2
We address the problem of stochastic combinatorial semi-bandits, where a player selects among P actions from the power set of a set containing d base items. Adaptivity to the probl…