On Commonsense Cues in BERT for Solving Commonsense Tasks
arXiv:2008.03945
Abstract
BERT has been used for solving commonsense tasks such as CommonsenseQA. While prior research has found that BERT does contain commonsense information to some extent, there has been work showing that pre-trained models can rely on spurious associations (e.g., data bias) rather than key cues in solving sentiment classification and other problems. We quantitatively investigate the presence of structural commonsense cues in BERT when solving commonsense tasks, and the importance of such cues for the model prediction. Using two different measures, we find that BERT does use relevant knowledge for solving the task, and the presence of commonsense knowledge is positively correlated to the model accuracy.
References in corpus (7)
- A Simple Method for Commonsense Reasoning
- Assessing BERT's Syntactic Abilities
- ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
- Do Attention Heads in BERT Track Syntactic Dependencies?
- Align, Mask and Select: A Simple Method for Incorporating Commonsense Knowledge into Language Representation Models
- Attention Interpretability Across NLP Tasks
- Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering