4 papers
Robust Markov Decision Processes on Continuous State Spaces
Mengmeng Li, Yifan Hu, Daniel Kuhn +1
We study infinite-horizon robust Markov decision processes (MDPs) on continuous state spaces with structured rectangular ambiguity set. The proposed ambiguity set falls within the…
Efficient Best-of-Both-Worlds Algorithms for Contextual Combinatorial Semi-Bandits
Mengmeng Li, Philipp J. Schneider, Jelisaveta AleksiÄ +1
We introduce the first best-of-both-worlds algorithm for contextual combinatorial semi-bandits that simultaneously guarantees regret in the adve…
Towards Optimal Offline Reinforcement Learning
Mengmeng Li, Daniel Kuhn, Tobias Sutter
We study offline reinforcement learning problems with a long-run average reward objective. The state-action pairs generated by any fixed behavioral policy thus follow a Markov chai…
Optimism in the Face of Ambiguity Principle for Multi-Armed Bandits
Mengmeng Li, Daniel Kuhn, Bahar TaÅkesen
Follow-The-Regularized-Leader (FTRL) algorithms often enjoy optimal regret for adversarial as well as stochastic bandit problems and allow for a streamlined analysis. Nonetheless,…