The Life Cycle of Knowledge in Big Language Models: A Survey
arXiv:2303.07616 · doi:10.1007/s11633-023-1416-x
Abstract
Knowledge plays a critical role in artificial intelligence. Recently, the extensive success of pre-trained language models (PLMs) has raised significant attention about how knowledge can be acquired, maintained, updated and used by language models. Despite the enormous amount of related studies, there still lacks a unified view of how knowledge circulates within language models throughout the learning, tuning, and application processes, which may prevent us from further understanding the connections between current progress or realizing existing limitations. In this survey, we revisit PLMs as knowledge-based systems by dividing the life circle of knowledge in PLMs into five critical periods, and investigating how knowledge circulates when it is built, maintained and used. To this end, we systematically review existing studies of each period of the knowledge life cycle, summarize the main challenges and current limitations, and discuss future directions.
paperlist: https://github.com/c-box/KnowledgeLifecycle
References in corpus (28)
- Training language models to follow instructions with human feedback
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- ERNIE: Enhanced Representation through Knowledge Integration
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- REALM: Retrieval-Augmented Language Model Pre-Training
- Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
- Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
- Assessing BERT's Syntactic Abilities
- SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability
- Avoiding Discrimination through Causal Reasoning
- What do you learn from context? Probing for sentence structure in contextualized word representations
- Transformers learn in-context by gradient descent
- Paradigm Shift in Natural Language Processing
- Generate rather than Retrieve: Large Language Models are Strong Context Generators
- Selective Annotation Makes Language Models Better Few-Shot Learners
- Data Distributional Properties Drive Emergent In-Context Learning in Transformers
- Explanations from Large Language Models Make Small Reasoners Better
- Can Large Language Models Truly Understand Prompts? A Case Study with Negated Prompts
- Retrieval-Augmented Multimodal Language Modeling
- A Survey of Knowledge-Intensive NLP with Pre-Trained Language Models
- Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language Model
- Do Language Models Have Beliefs? Methods for Detecting, Updating, and Visualizing Model Beliefs
- Honest Students from Untrusted Teachers: Learning an Interpretable Question-Answering Pipeline from a Pretrained Language Model
- Finding patterns in Knowledge Attribution for Transformers
- ThinkSum: Probabilistic reasoning over sets using large language models