7 citations · 12 across the 2 of their papers we have counts for
Showing 2020Show all
2 papers · 1 filter
cs.LG2020
Adapting to Delays and Data in Adversarial Multi-Armed Bandits
András György, Pooria Joulani
We consider the adversarial multi-armed bandit problem under delayed feedback. We analyze variants of the Exp3 algorithm that tune their step-size using only information (about the…
cs.LG2020
Adaptive Approximate Policy Iteration
Botao Hao, Nevena Lazic, Yasin Abbasi-Yadkori +2
Model-free reinforcement learning algorithms combined with value function approximation have recently achieved impressive performance in a variety of application domains. However,…