collaborators

5 papers

stat.ML2026

Efficient Best-of-Both-Worlds Algorithms for Contextual Combinatorial Semi-Bandits

Mengmeng Li, Philipp J. Schneider, Jelisaveta Aleksić +1

We introduce the first best-of-both-worlds algorithm for contextual combinatorial semi-bandits that simultaneously guarantees regret in the adve…

cs.LG2026

Soft-Radial Projection for Constrained End-to-End Learning

Philipp J. Schneider, Daniel Kuhn

Integrating hard constraints into deep learning is essential for safety-critical systems. Yet existing constructive layers that project predictions onto constraint boundaries face…

math.OC2025

Optimality of Linear Policies in Distributionally Robust Linear Quadratic Control

Bahar Taşkesen, Dan A. Iancu, Çağıl Koçyiğit +1

We study a generalization of the classical discrete-time, Linear-Quadratic-Gaussian (LQG) control problem where the noise distributions affecting the states and observations are un…

math.OC2025

Towards Optimal Offline Reinforcement Learning

Mengmeng Li, Daniel Kuhn, Tobias Sutter

We study offline reinforcement learning problems with a long-run average reward objective. The state-action pairs generated by any fixed behavioral policy thus follow a Markov chai…

cs.LG2025

Optimism in the Face of Ambiguity Principle for Multi-Armed Bandits

Mengmeng Li, Daniel Kuhn, Bahar Taşkesen

Follow-The-Regularized-Leader (FTRL) algorithms often enjoy optimal regret for adversarial as well as stochastic bandit problems and allow for a streamlined analysis. Nonetheless,…