3 papers
cs.LG2023
Delayed Bandits: When Do Intermediate Observations Help?
Emmanuel Esposito, Saeed Masoudian, Hao Qiu +3
We study a -armed bandit with delayed feedback and intermediate observations. We consider a model where intermediate observations have a form of a finite state, which is observe…
cs.LG2023
A Unified Analysis of Nonstochastic Delayed Feedback for Combinatorial Semi-Bandits, Linear Bandits, and MDPs
Dirk van der Hoeven, Lukas Zierahn, Tal Lancewicki +2
We derive a new analysis of Follow The Regularized Leader (FTRL) for online learning with delayed bandit feedback. By separating the cost of delayed feedback from that of bandit fe…
cs.LG2021
Nonstochastic Bandits and Experts with Arm-Dependent Delays
Dirk van der Hoeven, Nicolò Cesa-Bianchi
We study nonstochastic bandits and experts in a delayed setting where delays depend on both time and arms. While the setting in which delays only depend on time has been extensivel…