Just Sort It! A Simple and Effective Approach to Active Preference Learning
arXiv:1502.05556
Abstract
We address the problem of learning a ranking by using adaptively chosen pairwise comparisons. Our goal is to recover the ranking accurately but to sample the comparisons sparingly. If all comparison outcomes are consistent with the ranking, the optimal solution is to use an efficient sorting algorithm, such as Quicksort. But how do sorting algorithms behave if some comparison outcomes are inconsistent with the ranking? We give favorable guarantees for Quicksort for the popular Bradley-Terry model, under natural assumptions on the parameters. Furthermore, we empirically demonstrate that sorting algorithms lead to a very simple and effective active learning strategy: repeatedly sort the items. This strategy performs as well as state-of-the-art methods (and much better than random sampling) at a minuscule fraction of the computational cost.
Accepted at ICML 2017
Cited by in corpus (13)
- The Eighth Dialog System Technology Challenge
- A practical guide and software for analysing pairwise comparison experiments
- Preference-based Online Learning with Dueling Bandits: A Survey
- APRIL: Interactively Learning to Summarise by Combining Active Preference Learning and Reinforcement Learning
- Comparison Based Learning from Weak Oracles
- Hybrid Generative-Retrieval Transformers for Dialogue Domain Adaptation
- Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation
- Learning to Sample: an Active Learning Framework
- Graph Resistance and Learning from Pairwise Comparisons
- Robust Active Learning for Electrocardiographic Signal Classification
- Data-Efficient Methods for Dialogue Systems
- Let Me At Least Learn What You Really Like: Dealing With Noisy Humans When Learning Preferences
- Adaptive Sampling for Heterogeneous Rank Aggregation from Noisy Pairwise Comparisons