4 papers
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
Yao Zhao, Kwang-Sung Jun
Aligning large language models (LLMs) depends on high-quality datasets of human preference labels, which are costly to collect. Although active learning has been studied to improve…
Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors
Kapilan Balagopalan, Yinan Li, Yao Zhao +4
The best-arm identification (BAI) problem is one of the most fundamental problems in interactive machine learning, which has two flavors: the fixed-budget setting (FB) and the fixe…
Fixing the Loose Brake: Exponential-Tailed Stopping Time in Best Arm Identification
Kapilan Balagopalan, Tuan Ngo Nguyen, Yao Zhao +1
The best arm identification problem requires identifying the best alternative (i.e., arm) in active experimentation using the smallest number of experiments (i.e., arm pulls), whic…
Adaptive Experimentation When You Can't Experiment
Yao Zhao, Kwang-Sung Jun, Tanner Fiez +1
This paper introduces the \emph{confounded pure exploration transductive linear bandit} (\texttt{CPET-LB}) problem. As a motivating example, often online services cannot directly a…