1 paper
Pu Wang, Yao-Xiang Ding
We study N-armed stochastic dueling bandits under the Condorcet-winner assumption, where three widely adopted objectives are considered: best-arm identification (BAI), weak regre…