2 papers
cs.LG2025
A Tight Lower Bound for Non-stochastic Multi-armed Bandits with Expert Advice
Zachary Chase, Shinji Ito, Idan Mehalel
We determine the minimax optimal expected regret in the classic non-stochastic multi-armed bandit with expert advice problem, by proving a lower bound that matches the upper bound…
cs.LG2025
Deterministic Apple Tasting
Zachary Chase, Idan Mehalel
In binary () online classification with apple tasting feedback, the learner receives feedback only when predicting . Besides some degenerate learning tasks, all previously…