3 papers
cs.LG2026
Support Before Frequency in Discrete Diffusion
Adrian Müller, Antoine Gonon, Zebang Shen +2
Discrete diffusion models are increasingly competitive for language modeling, yet it remains unclear how their denoising objectives organize learning. Although these objectives tar…
cs.LG2025
Best of Both Worlds: Regret Minimization versus Minimax Play
Adrian Müller, Jon Schneider, Stratis Skoulakis +2
In this paper, we investigate the existence of online learning algorithms with bandit feedback that simultaneously guarantee regret compared to a given comparator strategy,…
cs.LG2024
Truly No-Regret Learning in Constrained MDPs
Adrian Müller, Pragnya Alatur, Volkan Cevher +2
Constrained Markov decision processes (CMDPs) are a common way to model safety constraints in reinforcement learning. State-of-the-art methods for efficiently solving CMDPs are bas…