Emergent Abilities of Large Language Models
arXiv:2206.07682
Abstract
Scaling up language models has been shown to predictably improve performance and sample efficiency on a wide range of downstream tasks. This paper instead discusses an unpredictable phenomenon that we refer to as emergent abilities of large language models. We consider an ability to be emergent if it is not present in smaller models but is present in larger models. Thus, emergent abilities cannot be predicted simply by extrapolating the performance of smaller models. The existence of such emergence implies that additional scaling could further expand the range of capabilities of language models.
Transactions on Machine Learning Research (TMLR), 2022
Cited by in corpus (77)
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- Opportunities and Challenges for ChatGPT and Large Language Models in Biomedicine and Health
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Bad Actor, Good Advisor: Exploring the Role of Large Language Models in Fake News Detection
- Evaluating Large Language Models on a Highly-specialized Topic, Radiation Oncology Physics
- Evaluation of Retrieval-Augmented Generation: A Survey
- ChatEDA: A Large Language Model Powered Autonomous Agent for EDA
- Evaluating the Effectiveness of LLMs in Introductory Computer Science Education: A Semester-Long Field Study
- ChartGPT: Leveraging LLMs to Generate Charts from Abstract Natural Language
- OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
- Materials science in the era of large language models: a perspective
- How understanding large language models can inform the use of ChatGPT in physics education
- Towards autonomous system: flexible modular production system enhanced with large language model agents
- LLM for SoC Security: A Paradigm Shift
- What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?
- Mathemyths: Leveraging Large Language Models to Teach Mathematical Language through Child-AI Co-Creative Storytelling
- General Purpose Artificial Intelligence Systems (GPAIS): Properties, Definition, Taxonomy, Societal Implications and Responsible Governance
- Deception Abilities Emerged in Large Language Models
- Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings
- KICGPT: Large Language Model with Knowledge in Context for Knowledge Graph Completion
- Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
- Learning from models beyond fine-tuning
- Identifying and Mitigating the Security Risks of Generative AI
- The Impact of ChatGPT and LLMs on Medical Imaging Stakeholders: Perspectives and Use Cases
- Towards an astronomical foundation model for stars with a Transformer-based model
- FoodSAM: Any Food Segmentation
- Achieving Peak Performance for Large Language Models: A Systematic Review
- Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review
- From Screens to Scenes: A Survey of Embodied AI in Healthcare
- Large language models as oracles for instantiating ontologies with domain-specific knowledge
- Prompted LLMs as Chatbot Modules for Long Open-domain Conversation
- Dual Use Concerns of Generative AI and Large Language Models
- Domain-specific ChatBots for Science using Embeddings
- Opportunities for Large Language Models and Discourse in Engineering Design
- Large Knowledge Model: Perspectives and Challenges
- Intelligent Computing Social Modeling and Methodological Innovations in Political Science in the Era of Large Language Models
- JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
- Comparing Traditional and LLM-based Search for Image Geolocation
- Efficient Transformers with Dynamic Token Pooling
- Computational Argumentation-based Chatbots: a Survey
- Solving Math Word Problems via Cooperative Reasoning induced Language Models
- LegiGPT: Party Politics and Transport Policy with Large Language Model
- Pre-Trained Language Models for Keyphrase Prediction: A Review
- What makes a language easy to deep-learn? Deep neural networks and humans similarly benefit from compositional structure
- Exploring the Landscape of Natural Language Processing Research
- Generative artificial intelligence in dentistry: Current approaches and future challenges
- Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
- Asymptotic theory of in-context learning by linear attention
- Large Language Models for Combinatorial Optimization: A Systematic Review
- Normative Conflicts and Shallow AI Alignment
- Judgment of Learning: A Human Ability Beyond Generative Artificial Intelligence
- Complex QA and language models hybrid architectures, Survey
- Arithmetic with Language Models: from Memorization to Computation
- Semi-supervised Multimodal Representation Learning through a Global Workspace
- Pre-Finetuning for Few-Shot Emotional Speech Recognition
- OPT-R: Exploring the Role of Explanations in Finetuning and Prompting for Reasoning Skills of Large Language Models
- GenComUI: Exploring Generative Visual Aids as Medium to Support Task-Oriented Human-Robot Communication
- Exploring the Impact of Model Scaling on Parameter-Efficient Tuning
- The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition
- Logic-Scaffolding: Personalized Aspect-Instructed Recommendation Explanation Generation using LLMs
- Automated Theorem Provers Help Improve Large Language Model Reasoning
- ESG Accountability Made Easy: DocQA at Your Service
- Assessing Logical Puzzle Solving in Large Language Models: Insights from a Minesweeper Case Study
- FlanEC: Exploring Flan-T5 for Post-ASR Error Correction
- Do GPT Language Models Suffer From Split Personality Disorder? The Advent Of Substrate-Free Psychometrics
- Using Artificial Populations to Study Psychological Phenomena in Neural Models
- Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
- A challenge in A(G)I, cybernetics revived in the Ouroboros Model as one algorithm for all thinking
- Assessing the nature of large language models: A caution against anthropocentrism
- Humanoid Artificial Consciousness Designed with Large Language Model Based on Psychoanalysis and Personality Theory
- Benchmarking quantized LLaMa-based models on the Brazilian Secondary School Exam
- Can "consciousness" be observed from large language model (LLM) internal states? Dissecting LLM representations obtained from Theory of Mind test with Integrated Information Theory and Span Representation analysis
- Emergent inabilities? Inverse scaling over the course of pretraining
- The LLM Mirage: Economic Interests and the Subversion of Weaponization Controls
- Continually Learn to Map Visual Concepts to Large Language Models in Resource-constrained Environments
- Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation
- Asking Better Questions -- The Art and Science of Forecasting: A mechanism for truer answers to high-stakes questions