Thresholding Bandit for Dose-ranging: The Impact of Monotonicity
arXiv:1711.04454
Abstract
We analyze the sample complexity of the thresholding bandit problem, with and without the assumption that the mean values of the arms are increasing. In each case, we provide a lower bound valid for any risk and any -correct algorithm; in addition, we propose an algorithm whose sample complexity is of the same order of magnitude for small risks. This work is motivated by phase 1 clinical trials, a practically important setting where the arm means are increasing by nature, and where no satisfactory solution is available so far.
Cited by in corpus (9)
- On Multi-Armed Bandit Designs for Dose-Finding Clinical Trials
- Gradient Ascent for Active Exploration in Bandit Problems
- Learning for Dose Allocation in Adaptive Clinical Trials with Safety Constraints
- Non-Asymptotic Pure Exploration by Solving Games
- SDF-Bayes: Cautious Optimism in Safe Dose-Finding Clinical Trials with Drug Combinations and Heterogeneous Patient Groups
- Best Arm Identification with Safety Constraints
- Learning with Comparison Feedback: Online Estimation of Sample Statistics
- The Influence of Shape Constraints on the Thresholding Bandit Problem
- Multi-armed Bandit Requiring Monotone Arm Sequences