Power Constrained Bandits
arXiv:2004.06230
Abstract
Contextual bandits often provide simple and effective personalization in decision making problems, making them popular tools to deliver personalized interventions in mobile health as well as other health applications. However, when bandits are deployed in the context of a scientific study -- e.g. a clinical trial to test if a mobile health intervention is effective -- the aim is not only to personalize for an individual, but also to determine, with sufficient statistical power, whether or not the system's intervention is effective. It is essential to assess the effectiveness of the intervention before broader deployment for better resource allocation. The two objectives are often deployed under different model assumptions, making it hard to determine how achieving the personalization and statistical power affect each other. In this work, we develop general meta-algorithms to modify existing algorithms such that sufficient power is guaranteed while still improving each user's well-being. We also demonstrate that our meta-algorithms are robust to various model mis-specifications possibly appearing in statistical studies, thus providing a valuable tool to study designers.
Accepted at MLHC 2021
References in corpus (3)
Cited by in corpus (8)
- Designing Reinforcement Learning Algorithms for Digital Interventions: Pre-implementation Guidelines
- Using Adaptive Bandit Experiments to Increase and Investigate Engagement in Mental Health
- Bandit Algorithms for Precision Medicine
- Online learning in bandits with predicted context
- Thompson sampling for zero-inflated count outcomes with an application to the Drink Less mobile health study
- Efficient Inference Without Trading-off Regret in Bandits: An Allocation Probability Test for Thompson Sampling
- Challenges in Statistical Analysis of Data Collected by a Bandit Algorithm: An Empirical Exploration in Applications to Adaptively Randomized Experiments
- Statistical Consequences of Dueling Bandits