GNN-LM: Language Modeling based on Global Contexts via GNN
arXiv:2110.08743
Abstract
Inspired by the notion that ``{\it to copy is easier than to memorize}``, in this work, we introduce GNN-LM, which extends the vanilla neural language model (LM) by allowing to reference similar contexts in the entire training corpus. We build a directed heterogeneous graph between an input context and its semantically related neighbors selected from the training corpus, where nodes are tokens in the input context and retrieved neighbor contexts, and edges represent connections between nodes. Graph neural networks (GNNs) are constructed upon the graph to aggregate information from similar contexts to decode the token. This learning paradigm provides direct access to the reference contexts and helps improve a model's generalization ability. We conduct comprehensive experiments to validate the effectiveness of the GNN-LM: GNN-LM achieves a new state-of-the-art perplexity of 14.8 on WikiText-103 (a 3.9 point improvement over its counterpart of the vanilla LM model), and shows substantial improvement on One Billion Word and Enwiki8 datasets against strong baselines. In-depth ablation studies are performed to understand the mechanics of GNN-LM. \footnote{The code can be found at https://github.com/ShannonAI/GNN-LM
To appear at ICLR 2022
References in corpus (22)
- Semi-Supervised Classification with Graph Convolutional Networks
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Longformer: The Long-Document Transformer
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- REALM: Retrieval-Augmented Language Model Pre-Training
- Pointer Sentinel Mixture Models
- Regularizing and Optimizing LSTM Language Models
- Pay Less Attention with Lightweight and Dynamic Convolutions
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval
- Data Noising as Smoothing in Neural Network Language Models
- Pre-training via Paraphrasing
- BP-Transformer: Modelling Long-Range Context via Binary Partitioning
- Generalization through Memorization: Nearest Neighbor Language Models
- BertGCN: Transductive Text Classification by Combining GCN and BERT
- Dynamic Evaluation of Transformer Language Models
- Efficient Retrieval Augmented Generation from Unstructured Knowledge for Task-Oriented Dialog
- Nearest Neighbor Machine Translation
- Augmenting Transformers with KNN-Based Composite Memory for Dialogue
- SAC: Accelerating and Structuring Self-Attention via Sparse Adaptive Connection
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
- Learning Dense Representations of Phrases at Scale
- Segatron: Segment-Aware Transformer for Language Modeling and Understanding