Measuring Word Significance using Distributed Representations of Words
arXiv:1508.02297
Abstract
Distributed representations of words as real-valued vectors in a relatively low-dimensional space aim at extracting syntactic and semantic features from large text corpora. A recently introduced neural network, named word2vec (Mikolov et al., 2013a; Mikolov et al., 2013b), was shown to encode semantic information in the direction of the word vectors. In this brief report, it is proposed to use the length of the vectors, together with the term frequency, as measure of word significance in a corpus. Experimental evidence using a domain-specific corpus of abstracts is presented to support this proposal. A useful visualization technique for text corpora emerges, where words are mapped onto a two-dimensional plane and automatically ranked by significance.
7 pages, 6 figures
Cited by in corpus (12)
- Meta-Path Guided Embedding for Similarity Search in Large-Scale Heterogeneous Information Networks
- Variational Graph Normalized Auto-Encoders
- Controlled Experiments for Word Embeddings
- Norm-Based Curriculum Learning for Neural Machine Translation
- Compass-aligned Distributional Embeddings for Studying Semantic Differences across Corpora
- Analyzing Structures in the Semantic Vector Space: A Framework for Decomposing Word Embeddings
- Retrofitting Vector Representations of Adverse Event Reporting Data to Structured Knowledge to Improve Pharmacovigilance Signal Detection
- Discrete Word Embedding for Logical Natural Language Understanding
- The presence of occupational structure in online texts based on word embedding NLP models
- Monitoring geometrical properties of word embeddings for detecting the emergence of new topics
- Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
- Learning Efficient Task-Specific Meta-Embeddings with Word Prisms