4 papers
Fixing the Loose Brake: Exponential-Tailed Stopping Time in Best Arm Identification
Kapilan Balagopalan, Tuan Ngo Nguyen, Yao Zhao +1
The best arm identification problem requires identifying the best alternative (i.e., arm) in active experimentation using the smallest number of experiments (i.e., arm pulls), whic…
HAVER: Instance-Dependent Error Bounds for Maximum Mean Estimation and Applications to Q-Learning and Monte Carlo Tree Search
Tuan Ngo Nguyen, Jay Barrett, Kwang-Sung Jun
We study the problem of estimating the \emph{value} of the largest mean among K distributions via samples from them (rather than estimating \emph{which} distribution has the larges…
Minimum Empirical Divergence for Sub-Gaussian Linear Bandits
Kapilan Balagopalan, Kwang-Sung Jun
We propose a novel linear bandit algorithm called LinMED (Linear Minimum Empirical Divergence), which is a linear extension of the MED algorithm that was originally designed for mu…
Adaptive Experimentation When You Can't Experiment
Yao Zhao, Kwang-Sung Jun, Tanner Fiez +1
This paper introduces the \emph{confounded pure exploration transductive linear bandit} (\texttt{CPET-LB}) problem. As a motivating example, often online services cannot directly a…