5 papers · 1 filter
Leveraging Similarities in Multi-Armed Bandits
Khaled Eldowa, Thibaud Rahier, Augustin Cablant +2
In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags, or hierarchical structure.…
Logarithmic Regret for Unconstrained Submodular Maximization Stochastic Bandit
Julien Zhou, Pierre Gaillard, Thibaud Rahier +1
We address the online unconstrained submodular maximization problem (Online USM), in a setting with stochastic bandit feedback. In this framework, a decision-maker receives noisy r…
Towards Efficient and Optimal Covariance-Adaptive Algorithms for Combinatorial Semi-Bandits
Julien Zhou, Pierre Gaillard, Thibaud Rahier +2
We address the problem of stochastic combinatorial semi-bandits, where a player selects among P actions from the power set of a set containing d base items. Adaptivity to the probl…
Zeroth-order non-convex learning via hierarchical dual averaging
Amélie Héliou, Matthieu Martin, Panayotis Mertikopoulos +1
We propose a hierarchical version of dual averaging for zeroth-order online non-convex optimization - i.e., learning processes where, at each stage, the optimizer is facing an unkn…
Online non-convex optimization with imperfect feedback
Amélie Héliou, Matthieu Martin, Panayotis Mertikopoulos +1
We consider the problem of online learning with non-convex losses. In terms of feedback, we assume that the learner observes - or otherwise constructs - an inexact model for the lo…