activity
20152020
most citedRegret Lower Bound and Optimal Algorithm in Dueling Bandit Problem

20 citations · 63 across the 9 of their papers we have counts for

collaborators
Showing stat.MLShow all

6 papers · 1 filter

stat.ML2020

Time-varying Gaussian Process Bandit Optimization with Non-constant Evaluation Time

Hideaki Imamura, Nontawat Charoenphakdee, Futoshi Futami +3

The Gaussian process bandit is a problem in which we want to find a maximizer of a black-box function with the minimum number of function evaluations. If the black-box function var…

stat.ML2019

On the Calibration of Multiclass Classification with Rejection

Chenri Ni, Nontawat Charoenphakdee, Junya Honda +1

We investigate the problem of multiclass classification with rejection, where a classifier can choose not to make a prediction to avoid critical misclassification. First, we consid…

stat.ML2018

Dueling Bandits with Qualitative Feedback

Liyuan Xu, Junya Honda, Masashi Sugiyama

We formulate and study a novel multi-armed bandit problem called the qualitative dueling bandit (QDB) problem, where an agent observes not numeric but qualitative feedback by pulli…

stat.ML201718 cited

Fully adaptive algorithm for pure exploration in linear bandits

Liyuan Xu, Junya Honda, Masashi Sugiyama

We propose the first fully-adaptive algorithm for pure exploration in linear bandits---the task to find the arm with the largest expected reward, which depends on an unknown parame…

stat.ML2016

Copeland Dueling Bandit Problem: Regret Lower Bound, Optimal Algorithm, and Computationally Efficient Algorithm

Junpei Komiyama, Junya Honda, Hiroshi Nakagawa

We study the K-armed dueling bandit problem, a variation of the standard stochastic bandit problem where the feedback is limited to relative comparisons of a pair of arms. The hard…

stat.ML201520 cited

Regret Lower Bound and Optimal Algorithm in Dueling Bandit Problem

Junpei Komiyama, Junya Honda, Hisashi Kashima +1

We study the -armed dueling bandit problem, a variation of the standard stochastic bandit problem where the feedback is limited to relative comparisons of a pair of arms. We int…