What you can cram into a single vector: Probing sentence embeddings for linguistic properties
arXiv:1805.01070
Abstract
Although much effort has recently been devoted to training high-quality sentence embeddings, we still have a poor understanding of what they are capturing. "Downstream" tasks, often based on sentence classification, are commonly used to evaluate the quality of sentence representations. The complexity of the tasks makes it however difficult to infer what kind of information is present in the representations. We introduce here 10 probing tasks designed to capture simple linguistic features of sentences, and we use them to study embeddings generated by three different encoders trained in eight distinct ways, uncovering intriguing properties of both encoders and training methods.
ACL 2018
Cited by in corpus (25)
- Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data Tasks
- Visualizing and Measuring the Geometry of BERT
- Do Vision Transformers See Like Convolutional Neural Networks?
- On the Effect of Dropping Layers of Pre-trained Transformer Models
- Correlating neural and symbolic representations of language
- Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans
- Language Modeling Teaches You More Syntax than Translation Does: Lessons Learned Through Auxiliary Task Analysis
- Multilingual NMT with a language-independent attention bridge
- Artificial Text Detection via Examining the Topology of Attention Maps
- Probing Neural Language Models for Human Tacit Assumptions
- Enhancing deep neural networks with morphological information
- Multi-Granularity Self-Attention for Neural Machine Translation
- Multi-Task Learning with Shared Encoder for Non-Autoregressive Machine Translation
- Self-Attention with Structural Position Representations
- Unsupervised Learning of Sentence Representations Using Sequence Consistency
- ShufText: A Simple Black Box Approach to Evaluate the Fragility of Text Classification Models
- Dissecting Contextual Word Embeddings: Architecture and Representation
- Assessing Phrasal Representation and Composition in Transformers
- The Rediscovery Hypothesis: Language Models Need to Meet Linguistics
- Learning Robust, Transferable Sentence Representations for Text Classification
- Towards Better Modeling Hierarchical Structure for Self-Attention with Ordered Neurons
- A Simple Recurrent Unit with Reduced Tensor Product Representations
- Picking Apart Story Salads
- COSTRA 1.0: A Dataset of Complex Sentence Transformations
- Exploiting Sentence Embedding for Medical Question Answering