5 papers
Provably Efficient Regularized Online RLHF with Generalized Bilinear Preferences
Junghyun Lee, Minju Hong, Kwang-Sung Jun +2
We consider the problem of regularized best-response max-regret minimization in online RLHF under general preferences and bandit feedback. While various regularizers are utilized t…
Fixed Budget is No Harder Than Fixed Confidence in Best-Arm Identification up to Logarithmic Factors
Kapilan Balagopalan, Yinan Li, Yao Zhao +4
The best-arm identification (BAI) problem is one of the most fundamental problems in interactive machine learning, which has two flavors: the fixed-budget setting (FB) and the fixe…
-Good Action Identification in Fixed-Budget Monte Carlo Tree Search
Yinan Li, Tuan Nguyen, Kwang-Sung Jun
We study the fixed-budget max-min action identification problem in depth-2 max-min trees, an important special case of Monte Carlo Tree Search. A learner sequentially allocates …
Second-Order Bounds for [0,1]-Valued Regression via Betting Loss
Yinan Li, Kwang-Sung Jun
We consider the -valued regression problem in the i.i.d. setting. In a related problem called cost-sensitive classification, \citet{foster21efficient} have shown that the lo…
Noise-Adaptive Confidence Sets for Linear Bandits and Application to Bayesian Optimization
Kwang-Sung Jun, Jungtaek Kim
Adapting to a priori unknown noise level is a very important but challenging problem in sequential decision-making as efficient exploration typically requires knowledge of the nois…