3 papers
cs.LG2025
Coherence Mechanisms for Provable Self-Improvement
Mehryar Mohri, Jon Schneider, Yifan Wu
Self-improvement is a critical capability for large language models and other intelligent systems, enabling them to refine their behavior and internal consistency without external…
cs.LG2025
Best of Both Worlds: Regret Minimization versus Minimax Play
Adrian Müller, Jon Schneider, Stratis Skoulakis +2
In this paper, we investigate the existence of online learning algorithms with bandit feedback that simultaneously guarantee regret compared to a given comparator strategy,…
cs.LG2025
A New Benchmark for Online Learning with Budget-Balancing Constraints
Mark Braverman, Jingyi Liu, Jieming Mao +2
The adversarial Bandit with Knapsack problem is a multi-armed bandits problem with budget constraints and adversarial rewards and costs. In each round, a learner selects an action…