3 papers
stat.ML2025
Efficient Best-of-Both-Worlds Algorithms for Contextual Combinatorial Semi-Bandits
Mengmeng Li, Philipp J. Schneider, Jelisaveta Aleksić +1
We introduce the first best-of-both-worlds algorithm for contextual combinatorial semi-bandits that simultaneously guarantees regret in the adve…
math.OC2025
Towards Optimal Offline Reinforcement Learning
Mengmeng Li, Daniel Kuhn, Tobias Sutter
We study offline reinforcement learning problems with a long-run average reward objective. The state-action pairs generated by any fixed behavioral policy thus follow a Markov chai…
cs.LG2024
Optimism in the Face of Ambiguity Principle for Multi-Armed Bandits
Mengmeng Li, Daniel Kuhn, Bahar Taşkesen
Follow-The-Regularized-Leader (FTRL) algorithms often enjoy optimal regret for adversarial as well as stochastic bandit problems and allow for a streamlined analysis. Nonetheless,…