REALM: Retrieval-Augmented Language Model Pre-Training
arXiv:2002.08909
The paper introduces REALM, a language model that retrieves relevant documents from a large corpus during pre‑training and inference, enabling it to use external knowledge for tasks like open‑domain question answering.
Abstract
Language model pre-training has been shown to capture a surprising amount of world knowledge, crucial for NLP tasks such as question answering. However, this knowledge is stored implicitly in the parameters of a neural network, requiring ever-larger networks to cover more facts. To capture knowledge in a more modular and interpretable way, we augment language model pre-training with a latent knowledge retriever, which allows the model to retrieve and attend over documents from a large corpus such as Wikipedia, used during pre-training, fine-tuning and inference. For the first time, we show how to pre-train such a knowledge retriever in an unsupervised manner, using masked language modeling as the learning signal and backpropagating through a retrieval step that considers millions of documents. We demonstrate the effectiveness of Retrieval-Augmented Language Model pre-training (REALM) by fine-tuning on the challenging task of Open-domain Question Answering (Open-QA). We compare against state-of-the-art models for both explicit and implicit knowledge storage on three popular Open-QA benchmarks, and find that we outperform all previous methods by a significant margin (4-16% absolute accuracy), while also providing qualitative benefits such as interpretability and modularity.
Topics & keywords
References in corpus (3)
Cited by in corpus (51)
- Language Models are Few-Shot Learners
- Long Range Arena: A Benchmark for Efficient Transformers
- Pre-training via Paraphrasing
- Multimodal Few-Shot Learning with Frozen Language Models
- How Context Affects Language Models' Factual Predictions
- RepBERT: Contextualized Text Embeddings for First-Stage Retrieval
- Machine Reading Comprehension: The Role of Contextualized Language Models and Beyond
- Modifying Memories in Transformer Models
- A Memory Efficient Baseline for Open Domain Question Answering
- Example-Based Named Entity Recognition
- Facts as Experts: Adaptable and Interpretable Neural Memory over Symbolic Knowledge
- Misinformation Has High Perplexity
- Embedding-based Zero-shot Retrieval through Query Generation
- End-to-End QA on COVID-19: Domain Adaptation with Synthetic Training
- Language Models As or For Knowledge Bases
- Neural Machine Translation with Monolingual Translation Memory
- Pre-Trained Models: Past, Present and Future
- Pre-training Text-to-Text Transformers for Concept-centric Common Sense
- SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval
- Optimizing Dense Retrieval Model Training with Hard Negatives
- Studying Strategically: Learning to Mask for Closed-book QA
- Retrieval-Augmented Transformer-XL for Close-Domain Dialog Generation
- Open-book Video Captioning with Retrieve-Copy-Generate Network
- One Question Answering Model for Many Languages with Cross-lingual Dense Passage Retrieval
- Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval
- Visual Grounding Strategies for Text-Only Natural Language Processing
- Good for Misconceived Reasons: An Empirical Revisiting on the Need for Visual Context in Multimodal Machine Translation
- PGT: Pseudo Relevance Feedback Using a Graph-Based Transformer
- Current Limitations of Language Models: What You Need is Retrieval
- Adaptive Semiparametric Language Models
- Training Large-Scale News Recommenders with Pretrained Language Models in the Loop
- Complementary Evidence Identification in Open-Domain Question Answering
- Towards Robust Neural Retrieval Models with Synthetic Pre-Training
- Augmented Abstractive Summarization With Document-LevelSemantic Graph
- SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking
- Neural Retrieval for Question Answering with Cross-Attention Supervised Data Augmentation
- Narrative Question Answering with Cutting-Edge Open-Domain QA Techniques: A Comprehensive Study
- Jointly Optimizing Query Encoder and Product Quantization to Improve Retrieval Performance
- Automatic Claim Review for Climate Science via Explanation Generation
- Frustratingly Hard Evidence Retrieval for QA Over Books
- Zero-shot Slot Filling with DPR and RAG
- Bew: Towards Answering Business-Entity-Related Web Questions
- Extreme Multi-label Learning for Semantic Matching in Product Search
- What makes us curious? analysis of a corpus of open-domain questions
- Hybrid Encoder: Towards Efficient and Precise Native AdsRecommendation via Hybrid Transformer Encoding Networks
- Instance-Based Neural Dependency Parsing
- Towards Universal Dense Retrieval for Open-domain Question Answering
- Relation-Guided Pre-Training for Open-Domain Question Answering
- Unsupervised Open-Domain Question Answering
- Coarse-to-Fine Memory Matching for Joint Retrieval and Classification
- Efficient Retrieval Optimized Multi-task Learning