11 citations · 13 across the 5 of their papers we have counts for
Showing 2020Show all
2 papers · 1 filter
cs.LG2020
Adaptive Algorithms for Multi-armed Bandit with Composite and Anonymous Feedback
Siwei Wang, Haoyun Wang, Longbo Huang
We study the multi-armed bandit (MAB) problem with composite and anonymous feedback. In this model, the reward of pulling an arm spreads over a period of time (we call this period…
cs.LG2020★ 11 cited
Restless-UCB, an Efficient and Low-complexity Algorithm for Online Restless Bandits
Siwei Wang, Longbo Huang, John C. S. Lui
We study the online restless bandit problem, where the state of each arm evolves according to a Markov chain, and the reward of pulling an arm depends on both the pulled arm and th…