From Paraphrase Database to Compositional Paraphrase Model and Back
arXiv:1506.03487
Abstract
The Paraphrase Database (PPDB; Ganitkevitch et al., 2013) is an extensive semantic resource, consisting of a list of phrase pairs with (heuristic) confidence estimates. However, it is still unclear how it can best be used, due to the heuristic nature of the confidences and its necessarily incomplete coverage. We propose models to leverage the phrase pairs from the PPDB to build parametric paraphrase models that score paraphrase pairs more accurately than the PPDB's internal scores while simultaneously improving its coverage. They allow for learning phrase embeddings as well as improved word embeddings. Moreover, we introduce two new, manually annotated datasets to evaluate short-phrase paraphrasing models. Using our paraphrase model trained using PPDB, we achieve state-of-the-art results on standard word and bigram similarity tasks and beat strong baselines on our new short phrase paraphrase tasks.
2015 TACL paper updated with an appendix describing new 300 dimensional embeddings. Submitted 1/2015. Accepted 2/2015. Published 6/2015
References in corpus (2)
Cited by in corpus (8)
- Unsupervised Learning of Sentence Embeddings using Compositional n-Gram Features
- Global-Locally Self-Attentive Dialogue State Tracker
- A Continuously Growing Dataset of Sentential Paraphrases
- Aff2Vec: Affect--Enriched Distributional Word Representations
- Paraphrase Detection on Noisy Subtitles in Six Languages
- P-SIF: Document Embeddings Using Partition Averaging
- One Representation per Word - Does it make Sense for Composition?
- Point or Generate Dialogue State Tracker