Efficient and Optimal Algorithms for Contextual Dueling Bandits under Realizability
arXiv:2111.12306
Abstract
We study the -armed contextual dueling bandit problem, a sequential decision making setting in which the learner uses contextual information to make two decisions, but only observes \emph{preference-based feedback} suggesting that one decision was better than the other. We focus on the regret minimization problem under realizability, where the feedback is generated by a pairwise preference matrix that is well-specified by a given function class . We provide a new algorithm that achieves the optimal regret rate for a new notion of best response regret, which is a strictly stronger performance measure than those considered in prior works. The algorithm is also computationally efficient, running in polynomial time assuming access to an online oracle for square loss regression over . This resolves an open problem of Dudík et al. [2015] on oracle efficient, regret-optimal algorithms for contextual dueling bandits.
References in corpus (11)
- Efficient Optimal Learning for Contextual Bandits
- Efficient Learning of Generalized Linear and Single Index Models with Isotonic Regression
- On the Universality of Online Mirror Descent
- Provable Self-Play Algorithms for Competitive Reinforcement Learning
- Contextual Bandit Learning with Predictable Rewards
- Reducing Dueling Bandits to Cardinal Bandits
- Instance-Dependent Complexity of Contextual Bandits and Reinforcement Learning: A Disagreement-Based Perspective
- Adapting to Misspecification in Contextual Bandits
- Optimal Dynamic Regret in Exp-Concave Online Learning
- Regret Analysis for Continuous Dueling Bandit
- Non-stationary Online Regression