4 papers · 1 filter
Leveraging Similarities in Multi-Armed Bandits
Khaled Eldowa, Thibaud Rahier, Augustin Cablant +2
In many online learning and bandit problems, the actions we consider possess inherent similarities--for instance because they share latent traits, tags, or hierarchical structure.…
Instance-Dependent Regret Bounds for Nonstochastic Linear Partial Monitoring
Federico Di Gennaro, Khaled Eldowa, Nicolò Cesa-Bianchi
In contrast to the classic formulation of partial monitoring, linear partial monitoring can model infinite outcome spaces, while imposing a linear structure on both the losses and…
Online Episodic Convex Reinforcement Learning
Bianca Marin Moreno, Khaled Eldowa, Pierre Gaillard +2
We study online learning in episodic finite-horizon Markov decision processes (MDPs) with convex objective functions, known as the concave utility reinforcement learning (CURL) pro…
Improved Regret Bounds for Bandits with Expert Advice
Nicolò Cesa-Bianchi, Khaled Eldowa, Emmanuel Esposito +1
In this research note, we revisit the bandits with expert advice problem. Under a restricted feedback model, we prove a lower bound of order for the worst-cas…