Ultradense Word Embeddings by Orthogonal Transformation
arXiv:1602.07572 · doi:10.18653/v1/N16-1091
Abstract
Embeddings are generic representations that are useful for many NLP tasks. In this paper, we introduce DENSIFIER, a method that learns an orthogonal transformation of the embedding space that focuses the information relevant for a task in an ultradense subspace of a dimensionality that is smaller by a factor of 100 than the original space. We show that ultradense embeddings generated by DENSIFIER reach state of the art on a lexicon creation task in which words are annotated with three types of lexical information - sentiment, concreteness and frequency. On the SemEval2015 10B sentiment analysis task we show that no information is lost when the ultradense subspace is used, but training is an order of magnitude more efficient due to the compactness of the ultradense space.
References in corpus (2)
Cited by in corpus (16)
- Discriminative Topic Mining via Category-Name Guided Text Embedding
- Predicting Concreteness and Imageability of Words Within and Across Languages via Word Embeddings
- Uncovering divergent linguistic information in word embeddings with lessons for intrinsic and extrinsic evaluation
- Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora
- Semi-Supervised Affective Meaning Lexicon Expansion Using Semantic and Distributed Word Representations
- Using Sentiment Induction to Understand Variation in Gendered Online Communities
- Second-Order Word Embeddings from Nearest Neighbor Topological Features
- Towards a science of human stories: using sentiment analysis and emotional arcs to understand the building blocks of complex social systems
- Learning Concept Abstractness Using Weak Supervision
- Expanding Subjective Lexicons for Social Media Mining with Embedding Subspaces
- Multi-lingual Common Semantic Space Construction via Cluster-consistent Word Embedding
- Analyzing Structures in the Semantic Vector Space: A Framework for Decomposing Word Embeddings
- Retrofitting Contextualized Word Embeddings with Paraphrases
- Detecting Domain Polarity-Changes of Words in a Sentiment Lexicon
- UniSent: Universal Adaptable Sentiment Lexica for 1000+ Languages
- Rotate King to get Queen: Word Relationships as Orthogonal Transformations in Embedding Space