Whittle index based Q-learning for restless bandits with average reward
arXiv:2004.14427 · doi:10.1016/j.automatica.2022.110186
Abstract
A novel reinforcement learning algorithm is introduced for multiarmed restless bandits with average reward, using the paradigms of Q-learning and Whittle index. Specifically, we leverage the structure of the Whittle index policy to reduce the search space of Q-learning, resulting in major computational gains. Rigorous convergence analysis is provided, supported by numerical experiments. The numerical experiments show excellent empirical performance of the proposed scheme.
References in corpus (1)
Cited by in corpus (11)
- User Dynamics-Aware Edge Caching and Computing for Mobile Virtual Reality
- Markovian restless bandits and index policies: A review
- Average-reward model-free reinforcement learning: a systematic review and literature mapping
- Testing Indexability and Computing Whittle and Gittins Index in Subcubic Time
- Learning Augmented Index Policy for Optimal Service Placement at the Network Edge
- Tabular and Deep Learning for the Whittle Index
- Whittle Index Based User Association in Dense Millimeter Wave Networks
- A Multi-Armed Bandit-based Approach to Mobile Network Provider Selection
- Detecting an Odd Restless Markov Arm with a Trembling Hand
- Lagrangian Index Policy for Restless Bandits with Average Reward
- Screening for an Infectious Disease as a Problem in Stochastic Control