5 papers
Efficient Best-of-Both-Worlds Algorithms for Contextual Combinatorial Semi-Bandits
Mengmeng Li, Philipp J. Schneider, Jelisaveta AleksiÄ +1
We introduce the first best-of-both-worlds algorithm for contextual combinatorial semi-bandits that simultaneously guarantees regret in the adve…
Soft-Radial Projection for Constrained End-to-End Learning
Philipp J. Schneider, Daniel Kuhn
Integrating hard constraints into deep learning is essential for safety-critical systems. Yet existing constructive layers that project predictions onto constraint boundaries face…
Optimality of Linear Policies in Distributionally Robust Linear Quadratic Control
Bahar TaÅkesen, Dan A. Iancu, ÃaÄıl KoçyiÄit +1
We study a generalization of the classical discrete-time, Linear-Quadratic-Gaussian (LQG) control problem where the noise distributions affecting the states and observations are un…
Towards Optimal Offline Reinforcement Learning
Mengmeng Li, Daniel Kuhn, Tobias Sutter
We study offline reinforcement learning problems with a long-run average reward objective. The state-action pairs generated by any fixed behavioral policy thus follow a Markov chai…
Optimism in the Face of Ambiguity Principle for Multi-Armed Bandits
Mengmeng Li, Daniel Kuhn, Bahar TaÅkesen
Follow-The-Regularized-Leader (FTRL) algorithms often enjoy optimal regret for adversarial as well as stochastic bandit problems and allow for a streamlined analysis. Nonetheless,…