How Reasonable are Common-Sense Reasoning Tasks: A Case-Study on the Winograd Schema Challenge and SWAG
arXiv:1811.01778 · doi:10.18653/v1/D19-1335
Abstract
Recent studies have significantly improved the state-of-the-art on common-sense reasoning (CSR) benchmarks like the Winograd Schema Challenge (WSC) and SWAG. The question we ask in this paper is whether improved performance on these benchmarks represents genuine progress towards common-sense-enabled systems. We make case studies of both benchmarks and design protocols that clarify and qualify the results of previous work by analyzing threats to the validity of previous experimental designs. Our protocols account for several properties prevalent in common-sense benchmarks including size limitations, structural regularities, and variable instance difficulty.
7 pages
References in corpus (4)
Cited by in corpus (13)
- Shortcut Learning in Deep Neural Networks
- A Surprisingly Robust Trick for Winograd Schema Challenge
- WinoGrande: An Adversarial Winograd Schema Challenge at Scale
- Align, Mask and Select: A Simple Method for Incorporating Commonsense Knowledge into Language Representation Models
- Recent Advances in Natural Language Inference: A Survey of Benchmarks, Resources, and Approaches
- Exploring Unsupervised Pretraining and Sentence Structure Modelling for Winograd Schema Challenge
- Beyond Leaderboards: A survey of methods for revealing weaknesses in Natural Language Inference data and models
- ASER: A Large-scale Eventuality Knowledge Graph
- Linguistically-Informed Transformations (LIT): A Method for Automatically Generating Contrast Sets
- Problems and Countermeasures in Natural Language Processing Evaluation
- Go Beyond Plain Fine-tuning: Improving Pretrained Models for Social Commonsense
- Attention-based Contrastive Learning for Winograd Schemas
- Towards Zero-shot Commonsense Reasoning with Self-supervised Refinement of Language Models