An Ensemble Method to Produce High-Quality Word Embeddings (2016)
arXiv:1604.01692
Abstract
A currently successful approach to computational semantics is to represent words as embeddings in a machine-learned vector space. We present an ensemble method that combines embeddings produced by GloVe (Pennington et al., 2014) and word2vec (Mikolov et al., 2013) with structured knowledge from the semantic networks ConceptNet (Speer and Havasi, 2012) and PPDB (Ganitkevitch et al., 2013), merging their information into a common representation with a large, multilingual vocabulary. The embeddings it produces achieve state-of-the-art performance on many word-similarity evaluations. Its score of on an evaluation of rare words (Luong et al., 2013) is 16% higher than the previous best known system.
Corrected author name, revised reproducibility instructions that didn't work anymore. 12 pages, 3 figures
References in corpus (1)
Cited by in corpus (19)
- Emotion Detection in Text: a Review
- Comparative Analysis of Word Embeddings for Capturing Word Similarities
- Enhance word representation for out-of-vocabulary on Ubuntu dialogue corpus
- Emotional Embeddings: Refining Word Embeddings to Capture Emotional Content of Words
- Evaluation of taxonomic and neural embedding methods for calculating semantic similarity
- Asynchronous Training of Word Embeddings for Large Text Corpora
- Story Cloze Ending Selection Baselines and Data Examination
- RETRO: Relation Retrofitting For In-Database Machine Learning on Textual Data
- Contextual Text Embeddings for Twi
- Learning Word Sense Embeddings from Word Sense Definitions
- Revisiting the Prepositional-Phrase Attachment Problem Using Explicit Commonsense Knowledge
- Multilingual Music Genre Embeddings for Effective Cross-Lingual Music Item Annotation
- Into the Battlefield: Quantifying and Modeling Intra-community Conflicts in Online Discussion
- Enhancing Semantic Word Representations by Embedding Deeper Word Relationships
- An Unsupervised Approach for Mapping between Vector Spaces
- Balancing the composition of word embeddings across heterogenous data sets
- Contributions to Representation Learning with Graph Autoencoders and Applications to Music Recommendation
- TAGLETS: A System for Automatic Semi-Supervised Learning with Auxiliary Data
- Wasserstein distances for evaluating cross-lingual embeddings