BERT Rediscovers the Classical NLP Pipeline
arXiv:1905.05950
Abstract
Pre-trained text encoders have rapidly advanced the state of the art on many NLP tasks. We focus on one such model, BERT, and aim to quantify where linguistic information is captured within the network. We find that the model represents the steps of the traditional NLP pipeline in an interpretable and localizable way, and that the regions responsible for each step appear in the expected sequence: POS tagging, parsing, NER, semantic roles, then coreference. Qualitative analysis reveals that the model can and often does adjust this pipeline dynamically, revising lower-level decisions on the basis of disambiguating information from higher-level representations.
Presented at ACL 2019
Cited by in corpus (19)
- A Benchmark Study of Machine Learning Models for Online Fake News Detection
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- Theoretical Limitations of Self-Attention in Neural Sequence Models
- Empirical Evaluation of Pre-trained Transformers for Human-Level NLP: The Role of Sample Size and Dimensionality
- ABNIRML: Analyzing the Behavior of Neural IR Models
- Latin BERT: A Contextual Language Model for Classical Philology
- A Benchmark for Lease Contract Review
- Towards Hierarchical Importance Attribution: Explaining Compositional Semantics for Neural Sequence Models
- The MultiBERTs: BERT Reproductions for Robustness Analysis
- HUBERT Untangles BERT to Improve Transfer across NLP Tasks
- Compositionality decomposed: how do neural networks generalise?
- Towards Evaluating the Robustness of Chinese BERT Classifiers
- Undivided Attention: Are Intermediate Layers Necessary for BERT?
- Does Vision-and-Language Pretraining Improve Lexical Grounding?
- BERT might be Overkill: A Tiny but Effective Biomedical Entity Linker based on Residual Convolutional Neural Networks
- Low-Complexity Probing via Finding Subnetworks
- Multi-Stream Transformers
- Can Edge Probing Tasks Reveal Linguistic Knowledge in QA Models?
- Semantically Driven Sentence Fusion: Modeling and Evaluation