Performance Impact Caused by Hidden Bias of Training Data for Recognizing Textual Entailment
arXiv:1804.08117
Abstract
The quality of training data is one of the crucial problems when a learning-centered approach is employed. This paper proposes a new method to investigate the quality of a large corpus designed for the recognizing textual entailment (RTE) task. The proposed method, which is inspired by a statistical hypothesis test, consists of two phases: the first phase is to introduce the predictability of textual entailment labels as a null hypothesis which is extremely unacceptable if a target corpus has no hidden bias, and the second phase is to test the null hypothesis using a Naive Bayes model. The experimental result of the Stanford Natural Language Inference (SNLI) corpus does not reject the null hypothesis. Therefore, it indicates that the SNLI corpus has a hidden bias which allows prediction of textual entailment labels from hypothesis sentences even if no context information is given by a premise sentence. This paper also presents the performance impact of NN models for RTE caused by this hidden bias.
Proceedings of the 11th International Conference on Language Resources and Evaluation (LREC2018)
Cited by in corpus (34)
- On the Opportunities and Risks of Foundation Models
- GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
- WinoGrande: An Adversarial Winograd Schema Challenge at Scale
- Adversarial NLI: A New Benchmark for Natural Language Understanding
- Explaining Question Answering Models through Text Generation
- Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing
- ANLIzing the Adversarial Natural Language Inference Dataset
- ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
- Asking Crowdworkers to Write Entailment Examples: The Best of Bad Options
- Analyzing Compositionality-Sensitivity of NLI Models
- DQI: Measuring Data Quality in NLP
- Can neural networks understand monotonicity reasoning?
- Do Neural Models Learn Systematicity of Monotonicity Inference in Natural Language?
- Counterfactually-Augmented SNLI Training Data Does Not Yield Better Generalization Than Unaugmented Data
- Diversify Your Datasets: Analyzing Generalization via Controlled Variance in Adversarial Datasets
- Our Evaluation Metric Needs an Update to Encourage Generalization
- Rissanen Data Analysis: Examining Dataset Characteristics via Description Length
- HELP: A Dataset for Identifying Shortcomings of Neural Models in Monotonicity Reasoning
- What do Deep Networks Like to Read?
- What is More Likely to Happen Next? Video-and-Language Future Event Prediction
- An Empirical Study on Model-agnostic Debiasing Strategies for Robust Natural Language Inference
- DocNLI: A Large-scale Dataset for Document-level Natural Language Inference
- The Curse of Performance Instability in Analysis Datasets: Consequences, Source, and Suggestions
- Masked Language Modeling and the Distributional Hypothesis: Order Word Matters Pre-training for Little
- SyGNS: A Systematic Generalization Testbed Based on Natural Language Semantics
- Denoising Large-Scale Image Captioning from Alt-text Data using Content Selection Models
- Mitigating Annotation Artifacts in Natural Language Inference Datasets to Improve Cross-dataset Generalization Ability
- Uncertain Natural Language Inference
- Exploring Transitivity in Neural NLI Models through Veridicality
- Robustness and Sensitivity of BERT Models Predicting Alzheimer's Disease from Text
- COM2SENSE: A Commonsense Reasoning Benchmark with Complementary Sentences
- A Logic-Based Framework for Natural Language Inference in Dutch
- To what extent do human explanations of model behavior align with actual model behavior?
- Avoiding Inference Heuristics in Few-shot Prompt-based Finetuning