Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
arXiv:1902.01007
Abstract
A machine learning system can score well on a given test set by relying on heuristics that are effective for frequent example types but break down in more challenging cases. We study this issue within natural language inference (NLI), the task of determining whether one sentence entails another. We hypothesize that statistical NLI models may adopt three fallible syntactic heuristics: the lexical overlap heuristic, the subsequence heuristic, and the constituent heuristic. To determine whether models have adopted these heuristics, we introduce a controlled evaluation set called HANS (Heuristic Analysis for NLI Systems), which contains many examples where the heuristics fail. We find that models trained on MNLI, including BERT, a state-of-the-art model, perform very poorly on HANS, suggesting that they have indeed adopted these heuristics. We conclude that there is substantial room for improvement in NLI systems, and that the HANS dataset can motivate and measure progress in this area
Camera-ready for ACL 2019
References in corpus (1)
Cited by in corpus (35)
- Language Models are Few-Shot Learners
- Shortcut Learning in Deep Neural Networks
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Measuring and Reducing Gendered Correlations in Pre-trained Models
- Gradient Starvation: A Learning Proclivity in Neural Networks
- ABNIRML: Analyzing the Behavior of Neural IR Models
- A Systematic Analysis of Morphological Content in BERT Models for Multiple Languages
- Explaining and Improving Model Behavior with k Nearest Neighbor Representations
- A Systematic Assessment of Syntactic Generalization in Neural Language Models
- MonaLog: a Lightweight System for Natural Language Inference Based on Monotonicity
- Neural Language Generation: Formulation, Methods, and Evaluation
- A New Dataset for Natural Language Inference from Code-mixed Conversations
- Factual Probing Is [MASK]: Learning vs. Learning to Recall
- Beyond Leaderboards: A survey of methods for revealing weaknesses in Natural Language Inference data and models
- Causal Inference in Natural Language Processing: Estimation, Prediction, Interpretation and Beyond
- Does Data Augmentation Improve Generalization in NLP?
- Automatic Construction of Evaluation Suites for Natural Language Generation Datasets
- Lacking the embedding of a word? Look it up into a traditional dictionary
- Data-Efficient Pretraining via Contrastive Self-Supervision
- COGS: A Compositional Generalization Challenge Based on Semantic Interpretation
- Counterfactual Variable Control for Robust and Interpretable Question Answering
- Towards Robustifying NLI Models Against Lexical Dataset Biases
- Grounding Representation Similarity with Statistical Testing
- SesameBERT: Attention for Anywhere
- How Does Adversarial Fine-Tuning Benefit BERT?
- Quantifying the Task-Specific Information in Text-Based Classifications
- An Empirical Comparison of Instance Attribution Methods for NLP
- Effective Batching for Recurrent Neural Network Grammars
- Geometry matters: Exploring language examples at the decision boundary
- Dependency-Based Neural Representations for Classifying Lines of Programs
- Character-level Representations Improve DRS-based Semantic Parsing Even in the Age of BERT
- A Logic-Based Framework for Natural Language Inference in Dutch
- Evaluating German Transformer Language Models with Syntactic Agreement Tests
- Discriminatively-Tuned Generative Classifiers for Robust Natural Language Inference