Improving language models by retrieving from trillions of tokens
arXiv:2112.04426
Abstract
We enhance auto-regressive language models by conditioning on document chunks retrieved from a large corpus, based on local similarity with preceding tokens. With a trillion token database, our Retrieval-Enhanced Transformer (RETRO) obtains comparable performance to GPT-3 and Jurassic-1 on the Pile, despite using 25 fewer parameters. After fine-tuning, RETRO performance translates to downstream knowledge-intensive tasks such as question answering. RETRO combines a frozen Bert retriever, a differentiable encoder and a chunked cross-attention mechanism to predict tokens based on an order of magnitude more data than what is typically consumed during training. We typically train RETRO from scratch, yet can also rapidly RETROfit pre-trained transformers with retrieval and still achieve good performance. Our work opens up new avenues for improving language models through explicit memory at unprecedented scale.
Fix incorrect reported numbers in Table 14
Cited by in corpus (17)
- Diffusion Models in Vision: A Survey
- Data Augmentation Approaches in Natural Language Processing: A Survey
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Delving into LLM-assisted writing in biomedical publications through excess vocabulary
- Woodpecker: Hallucination Correction for Multimodal Large Language Models
- Natural Language Processing Models That Automate Programming Will Transform Chemistry Research and Teaching
- ConceptEVA: Concept-Based Interactive Exploration and Customization of Document Summaries
- Learning to Automate Follow-up Question Generation using Process Knowledge for Depression Triage on Reddit Posts
- Prompt Perturbation in Retrieval-Augmented Generation based Large Language Models
- Improving Contrastive Learning of Sentence Embeddings with Case-Augmented Positives and Retrieved Negatives
- Deanthropomorphising NLP: Can a Language Model Be Conscious?
- LAMBDA: A Large Model Based Data Agent
- Complex QA and language models hybrid architectures, Survey
- End-to-End Open Vocabulary Keyword Search With Multilingual Neural Representations
- Variational Open-Domain Question Answering
- When AI reviews science: Can we trust the referee?
- PEFA: Parameter-Free Adapters for Large-scale Embedding-based Retrieval Models