The Benchmark Lottery
arXiv:2107.07002
Abstract
The world of empirical machine learning (ML) strongly relies on benchmarks in order to determine the relative effectiveness of different algorithms and methods. This paper proposes the notion of "a benchmark lottery" that describes the overall fragility of the ML benchmarking process. The benchmark lottery postulates that many factors, other than fundamental algorithmic superiority, may lead to a method being perceived as superior. On multiple benchmark setups that are prevalent in the ML community, we show that the relative performance of algorithms may be altered significantly simply by choosing different benchmark tasks, highlighting the fragility of the current paradigms and potential fallacious interpretation derived from benchmarking ML methods. Given that every benchmark makes a statement about what it perceives to be important, we argue that this might lead to biased progress in the community. We discuss the implications of the observed phenomena and provide recommendations on mitigating them using multiple machine learning domains and communities as use cases, including natural language processing, computer vision, information retrieval, recommender systems, and reinforcement learning.
References in corpus (20)
- Distilling the Knowledge in a Neural Network
- Language Models are Few-Shot Learners
- Emergence of Locomotion Behaviours in Rich Environments
- DeepMind Control Suite
- Massively Parallel Methods for Deep Reinforcement Learning
- Challenges of Real-World Reinforcement Learning
- Revisiting ResNets: Improved Training and Scaling Strategies
- The Evolved Transformer
- Long Range Arena: A Benchmark for Efficient Transformers
- A Large-scale Study of Representation Learning with the Visual Task Adaptation Benchmark
- Neural Episodic Control
- Are we done with ImageNet?
- The Ladder: A Reliable Leaderboard for Machine Learning Competitions
- The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics
- Online and Offline Reinforcement Learning by Planning with a Learned Model
- The advantages of multiple classes for reducing overfitting from test set reuse
- Transferring Inductive Biases through Knowledge Distillation
- Significant Improvements over the State of the Art? A Case Study of the MS MARCO Document Ranking Leaderboard
- How Robust are Model Rankings: A Leaderboard Customization Approach for Equitable Evaluation
- Rip van Winkle's Razor: A Simple Estimate of Overfit to Test Data