Assessing BERT's Syntactic Abilities
arXiv:1901.05287
Abstract
I assess the extent to which the recently introduced BERT model captures English syntactic phenomena, using (1) naturally-occurring subject-verb agreement stimuli; (2) "coloreless green ideas" subject-verb agreement stimuli, in which content words in natural sentences are randomly replaced with words sharing the same part-of-speech and inflection; and (3) manually crafted stimuli for subject-verb agreement and reflexive anaphora phenomena. The BERT model performs remarkably well on all cases.
Cited by in corpus (53)
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model
- Do Attention Heads in BERT Track Syntactic Dependencies?
- Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
- Intermediate-Task Transfer Learning with Pretrained Models for Natural Language Understanding: When and Why Does It Work?
- Machine Reading Comprehension: The Role of Contextualized Language Models and Beyond
- Universal Text Representation from BERT: An Empirical Study
- Is Multilingual BERT Fluent in Language Generation?
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar Induction
- What does BERT Learn from Multiple-Choice Reading Comprehension Datasets?
- A Systematic Analysis of Morphological Content in BERT Models for Multiple Languages
- A Systematic Assessment of Syntactic Generalization in Neural Language Models
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Does injecting linguistic structure into language models lead to better alignment with brain recordings?
- Open Sesame: Getting Inside BERT's Linguistic Knowledge
- Pre-Trained Models: Past, Present and Future
- On Explaining Your Explanations of BERT: An Empirical Study with Sequence Classification
- Can RNNs learn Recursive Nested Subject-Verb Agreements?
- Syntactic Data Augmentation Increases Robustness to Inference Heuristics
- Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases
- Transformers Generalize Linearly
- Improving AMR Parsing with Sequence-to-Sequence Pre-training
- Parsing as Pretraining
- Do Neural Models Learn Systematicity of Monotonicity Inference in Natural Language?
- Which *BERT? A Survey Organizing Contextualized Encoders
- Question Generation from Paragraphs: A Tale of Two Hierarchical Models
- Natural Language Inference in Context -- Investigating Contextual Reasoning over Long Texts
- Scalable Syntax-Aware Language Models Using Knowledge Distillation
- On the Robustness of Language Encoders against Grammatical Errors
- How Does Adversarial Fine-Tuning Benefit BERT?
- Lessons Learned from Applying off-the-shelf BERT: There is no Silver Bullet
- Improving BERT Pretraining with Syntactic Supervision
- On the Copying Behaviors of Pre-Training for Neural Machine Translation
- Intrinsic Knowledge Evaluation on Chinese Language Models
- Do Syntactic Probes Probe Syntax? Experiments with Jabberwocky Probing
- Weakly-Supervised Neural Response Selection from an Ensemble of Task-Specialised Dialogue Agents
- Specializing Word Embeddings (for Parsing) by Information Bottleneck
- IDS at SemEval-2020 Task 10: Does Pre-trained Language Model Know What to Emphasize?
- Are Transformers a Modern Version of ELIZA? Observations on French Object Verb Agreement
- A Relation-Oriented Clustering Method for Open Relation Extraction
- Frequency Effects on Syntactic Rule Learning in Transformers
- Improving Similar Language Translation With Transfer Learning
- Multi-Stream Transformers
- Learning Syntactic Dense Embedding with Correlation Graph for Automatic Readability Assessment
- Probing for Bridging Inference in Transformer Language Models
- Disentangling Semantics and Syntax in Sentence Embeddings with Pre-trained Language Models
- Sensing Ambiguity in Henry James' "The Turn of the Screw"
- Word Frequency Does Not Predict Grammatical Knowledge in Language Models
- A Generative Approach to Titling and Clustering Wikipedia Sections
- Causal Transformers Perform Below Chance on Recursive Nested Constructions, Unlike Humans
- Analysing the Effect of Masking Length Distribution of MLM: An Evaluation Framework and Case Study on Chinese MRC Datasets