Towards Universal Paraphrastic Sentence Embeddings
arXiv:1511.08198
Abstract
We consider the problem of learning general-purpose, paraphrastic sentence embeddings based on supervision from the Paraphrase Database (Ganitkevitch et al., 2013). We compare six compositional architectures, evaluating them on annotated textual similarity datasets drawn both from the same distribution as the training data and from a wide range of other domains. We find that the most complex architectures, such as long short-term memory (LSTM) recurrent neural networks, perform best on the in-domain data. However, in out-of-domain scenarios, simple architectures such as word averaging vastly outperform LSTMs. Our simplest averaging model is even competitive with systems tuned for the particular tasks while also being extremely efficient and easy to use. In order to better understand how these architectures compare, we conduct further experiments on three supervised NLP tasks: sentence similarity, entailment, and sentiment classification. We again find that the word averaging models perform well for sentence similarity and entailment, outperforming LSTMs. However, on sentiment classification, we find that the LSTM performs very strongly-even recording new state-of-the-art performance on the Stanford Sentiment Treebank. We then demonstrate how to combine our pretrained sentence embeddings with these supervised tasks, using them both as a prior and as a black box feature extractor. This leads to performance rivaling the state of the art on the SICK similarity and entailment tasks. We release all of our resources to the research community with the hope that they can serve as the new baseline for further work on universal sentence embeddings.
Published as a conference paper at ICLR 2016
References in corpus (14)
- Adam: A Method for Stochastic Optimization
- Sequence to Sequence Learning with Neural Networks
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- LSTM: A Search Space Odyssey
- Natural Language Processing (almost) from Scratch
- On the difficulty of training Recurrent Neural Networks
- Teaching Machines to Read and Comprehend
- Theano: new features and speed improvements
- Convolutional Neural Network Architectures for Matching Natural Language Sentences
- Grammar as a Foreign Language
- A Hierarchical Neural Autoencoder for Paragraphs and Documents
- Question Answering with Subgraph Embeddings
- Classifying Relations via Long Short Term Memory Networks along Shortest Dependency Path
- Representing Meaning with a Combination of Logical and Distributional Models
Cited by in corpus (39)
- Model-Agnostic Interpretability of Machine Learning
- How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation
- An efficient framework for learning sentence representations
- Learning to Generate Reviews and Discovering Sentiment
- SentEval: An Evaluation Toolkit for Universal Sentence Representations
- Neural Machine Translation Inspired Binary Code Similarity Comparison beyond Function Pairs
- Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
- End-to-End Retrieval in Continuous Space
- Machine Common Sense Concept Paper
- DeClarE: Debunking Fake News and False Claims using Evidence-Aware Deep Learning
- Combining Convolution and Recursive Neural Networks for Sentiment Analysis
- Word Mover's Embedding: From Word2Vec to Document Embedding
- Effective Representations of Clinical Notes
- Learning Sentence Representation with Guidance of Human Attention
- Better Text Understanding Through Image-To-Text Transfer
- Multi-Label Transfer Learning for Multi-Relational Semantic Similarity
- ADSAGE: Anomaly Detection in Sequences of Attributed Graph Edges applied to insider threat detection at fine-grained level
- Causal Explanation Analysis on Social Media
- Representing Sentences as Low-Rank Subspaces
- How To Evaluate Your Dialogue System: Probe Tasks as an Alternative for Token-level Evaluation Metrics
- How neural networks learn to classify chaotic time series
- Fake Sentence Detection as a Training Task for Sentence Encoding
- An Iterative Polishing Framework based on Quality Aware Masked Language Model for Chinese Poetry Generation
- Syntax-based Attention Model for Natural Language Inference
- Assessing Dialogue Systems with Distribution Distances
- Latent Semantic Analysis Approach for Document Summarization Based on Word Embeddings
- Direct Network Transfer: Transfer Learning of Sentence Embeddings for Semantic Similarity
- Discovery of Latent Factors in High-dimensional Data Using Tensor Methods
- Leveraging Lexical Resources for Learning Entity Embeddings in Multi-Relational Data
- Large Scale Product Categorization using Structured and Unstructured Attributes
- Exploring Sentence Vector Spaces through Automatic Summarization
- Exploiting Review Neighbors for Contextualized Helpfulness Prediction
- SentPWNet: A Unified Sentence Pair Weighting Network for Task-specific Sentence Embedding
- Dynamic Embedding on Textual Networks via a Gaussian Process
- Learning to Learn Semantic Parsers from Natural Language Supervision
- Variational learning across domains with triplet information
- Paraphrase Thought: Sentence Embedding Module Imitating Human Language Recognition
- Hamming Sentence Embeddings for Information Retrieval
- Semantic maps and metrics for science Semantic maps and metrics for science using deep transformer encoders