FRAGE: Frequency-Agnostic Word Representation
arXiv:1809.06858
Abstract
Continuous word representation (aka word embedding) is a basic building block in many neural network-based models used in natural language processing tasks. Although it is widely accepted that words with similar semantics should be close to each other in the embedding space, we find that word embeddings learned in several tasks are biased towards word frequency: the embeddings of high-frequency and low-frequency words lie in different subregions of the embedding space, and the embedding of a rare word and a popular word can be far from each other even if they are semantically similar. This makes learned word embeddings ineffective, especially for rare words, and consequently limits the performance of these neural network models. In this paper, we develop a neat, simple yet effective way to learn \emph{FRequency-AGnostic word Embedding} (FRAGE) using adversarial training. We conducted comprehensive studies on ten datasets across four natural language processing tasks, including word similarity, language modeling, machine translation and text classification. Results show that with FRAGE, we achieve higher performance than the baselines in all tasks.
To appear in NIPS 2018
Cited by in corpus (23)
- AutoML: A Survey of the State-of-the-Art
- MaxUp: A Simple Way to Improve Generalization of Neural Network Training
- HerBERT: Efficiently Pretrained Transformer-based Language Model for Polish
- Calibration, Entropy Rates, and Memory in Language Models
- Understanding Recurrent Neural Architectures by Analyzing and Synthesizing Long Distance Dependencies in Benchmark Sequential Datasets
- Double-Hard Debias: Tailoring Word Embeddings for Gender Bias Mitigation
- Taking Notes on the Fly Helps BERT Pre-training
- Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive Translation
- RNNs Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?
- Automated Source Code Generation and Auto-completion Using Deep Learning: Comparing and Discussing Current Language-Model-Related Approaches
- ElixirNet: Relation-aware Network Architecture Adaptation for Medical Lesion Detection
- MDR Cluster-Debias: A Nonlinear WordEmbedding Debiasing Pipeline
- Multimodal Embeddings from Language Models
- Learning to Remove: Towards Isotropic Pre-trained BERT Embedding
- Telephonetic: Making Neural Language Models Robust to ASR and Semantic Noise
- A Simple Recurrent Unit with Reduced Tensor Product Representations
- Out-of-Manifold Regularization in Contextual Embedding Space for Text Classification
- SimpleBooks: Long-term dependency book dataset with simplified English vocabulary for word-level language modeling
- Neural Language Modeling With Implicit Cache Pointers
- Recoding latent sentence representations -- Dynamic gradient-based activation modification in RNNs
- Low Frequency Names Exhibit Bias and Overfitting in Contextualizing Language Models
- Barack's Wife Hillary: Using Knowledge-Graphs for Fact-Aware Language Modeling
- Adaptive Noise Injection: A Structure-Expanding Regularization for RNN