Which *BERT? A Survey Organizing Contextualized Encoders
arXiv:2010.00854
Abstract
Pretrained contextualized text encoders are now a staple of the NLP community. We present a survey on language representation learning with the aim of consolidating a series of shared lessons learned across a variety of recent efforts. While significant advancements continue at a rapid pace, we find that enough has now been discovered, in different directions, that we can begin to organize advances according to common themes. Through this organization, we highlight important considerations when interpreting recent contributions and choosing which model to use.
EMNLP 2020
References in corpus (13)
- Distilling the Knowledge in a Neural Network
- Scaling Laws for Neural Language Models
- ERNIE: Enhanced Representation through Knowledge Integration
- Multilingual Denoising Pre-training for Neural Machine Translation
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- Generating Long Sequences with Sparse Transformers
- Reformer: The Efficient Transformer
- Assessing BERT's Syntactic Abilities
- Cross-Lingual Ability of Multilingual BERT: An Empirical Study
- Insertion Transformer: Flexible Sequence Generation via Insertion Operations
- Multilingual Alignment of Contextual Word Representations
- KERMIT: Generative Insertion-Based Modeling for Sequences
- BP-Transformer: Modelling Long-Range Context via Binary Partitioning