Word Emdeddings through Hellinger PCA
arXiv:1312.5542
Abstract
Word embeddings resulting from neural language models have been shown to be successful for a large variety of NLP tasks. However, such architecture might be difficult to train and time-consuming. Instead, we propose to drastically simplify the word embeddings computation through a Hellinger PCA of the word co-occurence matrix. We compare those new word embeddings with some well-known embeddings on NER and movie review tasks and show that we can reach similar or even better performance. Although deep learning is not really necessary for generating good word embeddings, we show that it can provide an easy way to adapt embeddings to specific tasks.
9 pages, 5 tables
References in corpus (2)
Cited by in corpus (9)
- Recurrent neural networks with specialized word embeddings for health-domain named-entity recognition
- Neural Models for Information Retrieval
- Bidirectional LSTM-CRF for Clinical Concept Extraction
- Word Embedding based New Corpus for Low-resourced Language: Sindhi
- Attention Fusion Networks: Combining Behavior and E-mail Content to Improve Customer Support
- Sparse Lifting of Dense Vectors: Unifying Word and Sentence Representations
- Fast Label Embeddings via Randomized Linear Algebra
- Enhancing Semantic Word Representations by Embedding Deeper Word Relationships
- Meta-Embedding as Auxiliary Task Regularization