2 papers
cs.LG2026
Learning Safely Without Knowing the World:COMPASS-Hedge
Ting Hu, Luanda Cai, Emmanouil-Vasileios Vlatakis-Gkaragkounis
Online learning algorithms often face a fundamental trilemma: balancing regret guarantees between adversarial and stochastic settings and providing baseline safety against a fixed…
cs.LG2026
Prudent-Banker: No Extra Fees for Baseline Safety in Adversarial Bandits With and Without Delays
Ting Hu, Luanda Cai, Emmanouil-Vasileios Vlatakis-Gkaragkounis
We study adversarial multi-armed bandits with and without delayed feedback under a safety-aware goal: achieving minimax-optimal worst-case regret while keeping nearly constant regr…