paper

Optimal Thresholding Linear Bandit

arXiv:2402.09467

Abstract

We study a novel pure exploration problem: the -Thresholding Bandit Problem (TBP) with fixed confidence in stochastic linear bandits. We prove a lower bound for the sample complexity and extend an algorithm designed for Best Arm Identification in the linear case to TBP that is asymptotically optimal.

arXiv admin note: substantial text overlap with arXiv:2006.16073 by other authors