On the Dimensionality of Word Embedding
arXiv:1812.04224
Abstract
In this paper, we provide a theoretical understanding of word embedding and its dimensionality. Motivated by the unitary-invariance of word embedding, we propose the Pairwise Inner Product (PIP) loss, a novel metric on the dissimilarity between word embeddings. Using techniques from matrix perturbation theory, we reveal a fundamental bias-variance trade-off in dimensionality selection for word embeddings. This bias-variance trade-off sheds light on many empirical observations which were previously unexplained, for example the existence of an optimal dimensionality. Moreover, new insights and discoveries, like when and how word embeddings are robust to over-fitting, are revealed. By optimizing over the bias-variance trade-off of the PIP loss, we can explicitly answer the open question of dimensionality selection for word embedding.
18 pages, Advances in Neural Information Processing Systems 31 (NeurIPS 2018, Oral Presentation)
References in corpus (3)
Cited by in corpus (17)
- DDGK: Learning Graph Representations for Deep Divergence Graph Kernels
- Temporal Aggregation and Propagation Graph Neural Networks for Dynamic Representation
- On the Dimensionality of Embeddings for Sparse Features and Data
- Sequence embeddings help to identify fraudulent cases in healthcare insurance
- Evaluation of taxonomic and neural embedding methods for calculating semantic similarity
- Training with Multi-Layer Embeddings for Model Reduction
- Word2vec Skip-gram Dimensionality Selection via Sequential Normalized Maximum Likelihood
- Effects of Number of Filters of Convolutional Layers on Speech Recognition Model Accuracy
- Task-Feature Collaborative Learning with Application to Personalized Attribute Prediction
- Sparse Lifting of Dense Vectors: Unifying Word and Sentence Representations
- Decompressing Knowledge Graph Representations for Link Prediction
- Blind signal decomposition of various word embeddings based on join and individual variance explained
- Inflating Topic Relevance with Ideology: A Case Study of Political Ideology Bias in Social Topic Detection Models
- Learning Word Embeddings with Domain Awareness
- Word Embeddings: Stability and Semantic Change
- RPD: A Distance Function Between Word Embeddings
- SEMIE: SEMantically Infused Embeddings with Enhanced Interpretability for Domain-specific Small Corpus