A Mutual Information Maximization Perspective of Language Representation Learning
arXiv:1910.08350
Abstract
We show state-of-the-art word representation learning methods maximize an objective function that is a lower bound on the mutual information between different parts of a word sequence (i.e., a sentence). Our formulation provides an alternative perspective that unifies classical word embedding models (e.g., Skip-gram) and modern contextual embeddings (e.g., BERT, XLNet). In addition to enhancing our theoretical understanding of these methods, our derivation leads to a principled framework that can be used to construct new self-supervised tasks. We provide an example by drawing inspirations from related methods based on mutual information maximization that have been successful in computer vision, and introduce a simple self-supervised objective that maximizes the mutual information between a global sentence representation and n-grams in the sentence. Our analysis offers a holistic view of representation learning methods to transfer knowledge and translate progress across multiple domains (e.g., natural language processing, computer vision, audio processing).
12 pages, 3 figures
References in corpus (5)
- Cross-lingual Language Model Pretraining
- Learning Representations by Maximizing Mutual Information Across Views
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- A Theoretical Analysis of Contrastive Unsupervised Representation Learning
- Learning and Evaluating General Linguistic Intelligence
Cited by in corpus (12)
- Contrastive Representation Learning: A Framework and Review
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing
- On Mutual Information in Contrastive Learning for Visual Representations
- The Effect of Natural Distribution Shift on Question Answering Models
- A Survey on Self-supervised Pre-training for Sequential Transfer Learning in Neural Networks
- Understand and Improve Contrastive Learning Methods for Visual Representation: A Review
- Evolution Is All You Need: Phylogenetic Augmentation for Contrastive Learning
- Improving Robustness and Generality of NLP Models Using Disentangled Representations
- Hybrid Generative-Contrastive Representation Learning
- Self-supervised Representation Learning with Relative Predictive Coding
- Self-Supervised Graph Learning with Proximity-based Views and Channel Contrast
- Cross-modal Image Retrieval with Deep Mutual Information Maximization