No Training Required: Exploring Random Encoders for Sentence Classification
arXiv:1901.10444
Abstract
We explore various methods for computing sentence representations from pre-trained word embeddings without any training, i.e., using nothing but random parameterizations. Our aim is to put sentence embeddings on more solid footing by 1) looking at how much modern sentence embeddings gain over random methods---as it turns out, surprisingly little; and by 2) providing the field with more appropriate baselines going forward---which are, as it turns out, quite strong. We also make important observations about proper experimental protocol for sentence classification evaluation, together with recommendations for future research.
Published as a conference paper at ICLR 2019
Cited by in corpus (21)
- Do Attention Heads in BERT Track Syntactic Dependencies?
- Pre-training via Paraphrasing
- Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text
- Efficient and Private Federated Learning with Partially Trainable Networks
- Learning Compressed Sentence Representations for On-Device Text Processing
- Echo State Neural Machine Translation
- On the Interpretability of Deep Learning Based Models for Knowledge Tracing
- Contextual Lensing of Universal Sentence Representations
- Neural Language Priors
- CBOW Is Not All You Need: Combining CBOW with the Compositional Matrix Space Model
- On the impressive performance of randomly weighted encoders in summarization tasks
- Unsupervised Natural Question Answering with a Small Model
- Sentence Embeddings by Ensemble Distillation
- Understanding Pure Character-Based Neural Machine Translation: The Case of Translating Finnish into English
- Cost-effective Deployment of BERT Models in Serverless Environment
- Gating Mechanisms for Combining Character and Word-level Word Representations: An Empirical Study
- Text Classification and Clustering with Annealing Soft Nearest Neighbor Loss
- COM2SENSE: A Commonsense Reasoning Benchmark with Complementary Sentences
- Text Classification with Lexicon from PreAttention Mechanism
- A Revised Generative Evaluation of Visual Dialogue
- Deep Contextualized Self-training for Low Resource Dependency Parsing