All-but-the-Top: Simple and Effective Postprocessing for Word Representations
arXiv:1702.01417
Abstract
Real-valued word representations have transformed NLP applications; popular examples are word2vec and GloVe, recognized for their ability to capture linguistic regularities. In this paper, we demonstrate a {\em very simple}, and yet counter-intuitive, postprocessing technique -- eliminate the common mean vector and a few top dominating directions from the word vectors -- that renders off-the-shelf representations {\em even stronger}. The postprocessing is empirically validated on a variety of lexical-level intrinsic tasks (word similarity, concept categorization, word analogy) and sentence-level tasks (semantic textural similarity and { text classification}) on multiple datasets and with a variety of representation methods and hyperparameter choices in multiple languages; in each case, the processed representations are consistently better than the original ones.
Cited by in corpus (9)
- Extracting Sentence Embeddings from Pretrained Transformer Models
- On the Sentence Embeddings from Pre-trained Language Models
- IsoScore: Measuring the Uniformity of Embedding Space Utilization
- An Empirical Study on Post-processing Methods for Word Embeddings
- Few-Shot Text Classification with Pre-Trained Word Embeddings and a Human in the Loop
- A Meta-embedding-based Ensemble Approach for ICD Coding Prediction
- Multi-view Sentence Representation Learning
- Improving information retrieval through correspondence analysis instead of latent semantic analysis
- A Comparative Study on Structural and Semantic Properties of Sentence Embeddings