Neural Language Priors
arXiv:1910.03492
Abstract
The choice of sentence encoder architecture reflects assumptions about how a sentence's meaning is composed from its constituent words. We examine the contribution of these architectures by holding them randomly initialised and fixed, effectively treating them as as hand-crafted language priors, and evaluating the resulting sentence encoders on downstream language tasks. We find that even when encoders are presented with additional information that can be used to solve tasks, the corresponding priors do not leverage this information, except in an isolated case. We also find that apparently uninformative priors are just as good as seemingly informative priors on almost all tasks, indicating that learning is a necessary component to leverage information provided by architecture choice.
4 pages, 1 figure, 1 table
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Understanding deep learning requires rethinking generalization
- fairseq: A Fast, Extensible Toolkit for Sequence Modeling
- Weighted Transformer Network for Machine Translation
- No Training Required: Exploring Random Encoders for Sentence Classification
- Bidirectional Tree-Structured LSTM with Head Lexicalization