Sparks of Artificial General Intelligence: Early experiments with GPT-4
arXiv:2303.12712
Abstract
Artificial intelligence (AI) researchers have been developing and refining large language models (LLMs) that exhibit remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. The latest model developed by OpenAI, GPT-4, was trained using an unprecedented scale of compute and data. In this paper, we report on our investigation of an early version of GPT-4, when it was still in active development by OpenAI. We contend that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models. We discuss the rising capabilities and implications of these models. We demonstrate that, beyond its mastery of language, GPT-4 can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more, without needing any special prompting. Moreover, in all of these tasks, GPT-4's performance is strikingly close to human-level performance, and often vastly surpasses prior models such as ChatGPT. Given the breadth and depth of GPT-4's capabilities, we believe that it could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system. In our exploration of GPT-4, we put special emphasis on discovering its limitations, and we discuss the challenges ahead for advancing towards deeper and more comprehensive versions of AGI, including the possible need for pursuing a new paradigm that moves beyond next-word prediction. We conclude with reflections on societal influences of the recent technological leap and future research directions.
Cited by in corpus (94)
- Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization
- ChatGPT Chemistry Assistant for Text Mining and Prediction of MOF Synthesis
- Graph of Thoughts: Solving Elaborate Problems with Large Language Models
- Differentiable modeling to unify machine learning and physical models and advance Geosciences
- Unleashing the potential of prompt engineering for large language models
- ChatGPT for Shaping the Future of Dentistry: The Potential of Multi-Modal Large Language Model
- Towards Human-centered Explainable AI: A Survey of User Studies for Model Explanations
- 14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon
- Game of Tones: Faculty detection of GPT-4 generated content in university assessments
- A GPT-4 Reticular Chemist for Guiding MOF Discovery
- Evaluating Large Language Models on a Highly-specialized Topic, Radiation Oncology Physics
- Sensecape: Enabling Multilevel Exploration and Sensemaking with Large Language Models
- Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses
- Large Language Models as Zero-Shot Conversational Recommenders
- On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial
- How understanding large language models can inform the use of ChatGPT in physics education
- General Purpose Artificial Intelligence Systems (GPAIS): Properties, Definition, Taxonomy, Societal Implications and Responsible Governance
- From task structures to world models: What do LLMs know?
- Deception Abilities Emerged in Large Language Models
- GPT has become financially literate: Insights from financial literacy tests of GPT and a preliminary test of how people use it as a source of advice
- Theory of Mind for Multi-Agent Collaboration via Large Language Models
- A newcomer's guide to deep learning for inverse design in nano-photonics
- The effect of source disclosure on evaluation of AI-generated messages: A two-part study
- Educational impacts of generative artificial intelligence on learning and performance of engineering students in China
- Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
- Exploring the psychology of LLMs' Moral and Legal Reasoning
- Generative AI and Process Systems Engineering: The Next Frontier
- A Comprehensive Survey on Integrating Large Language Models with Knowledge-Based Methods
- LLM-assisted Knowledge Graph Engineering: Experiments with ChatGPT
- Playing repeated games with Large Language Models
- HPC-GPT: Integrating Large Language Model for High-Performance Computing
- Generative Discovery of Novel Chemical Designs using Diffusion Modeling and Transformer Deep Neural Networks with Application to Deep Eutectic Solvents
- The Impact of ChatGPT and LLMs on Medical Imaging Stakeholders: Perspectives and Use Cases
- Towards an astronomical foundation model for stars with a Transformer-based model
- FoodSAM: Any Food Segmentation
- Selenite: Scaffolding Online Sensemaking with Comprehensive Overviews Elicited from Large Language Models
- Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review
- CoverUp: Effective High Coverage Test Generation for Python
- Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate
- Large-Scale Text Analysis Using Generative Language Models: A Case Study in Discovering Public Value Expressions in AI Patents
- Towards Reliable Generative AI-Driven Scaffolding: Reducing Hallucinations and Enhancing Quality in Self-Regulated Learning Support
- Domain-specific ChatBots for Science using Embeddings
- Supporting Energy Policy Research with Large Language Models
- Drastic Circuit Depth Reductions with Preserved Adversarial Robustness by Approximate Encoding for Quantum Machine Learning
- Large language models and linguistic intentionality
- Theory of Mind abilities of Large Language Models in Human-Robot Interaction : An Illusion?
- Standards for Belief Representations in LLMs
- GPT-assisted learning of structure-property relationships by graph neural networks: Application to rare-earth doped phosphors
- Computational Argumentation-based Chatbots: a Survey
- Large Language Models at Work in China's Labor Market
- Artificial intelligence is algorithmic mimicry: why artificial "agents" are not (and won't be) proper agents
- Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field
- Introduction to dynamical mean-field theory of randomly connected neural networks with bidirectionally correlated couplings
- Deep learning for nano-photonic materials -- The solution to everything!?
- The Science Fiction Science Method
- From Bytes to Biases: Investigating the Cultural Self-Perception of Large Language Models
- The human biological advantage over AI
- Transforming Agency. On the mode of existence of Large Language Models
- Leveraging Lecture Content for Improved Feedback: Explorations with GPT-4 and Retrieval Augmented Generation
- Symbolic Equation Solving via Reinforcement Learning
- Investigating the Day-to-Day Experiences of Users with Traumatic Brain Injury with Conversational Agents
- Large Language Models for Combinatorial Optimization: A Systematic Review
- What Does Success Look Like? Catalyzing Meeting Intentionality with AI-Assisted Prospective Reflection
- Relational Programming with Foundation Models
- Arithmetic with Language Models: from Memorization to Computation
- Unmasking the giant: A comprehensive evaluation of ChatGPT's proficiency in coding algorithms and data structures
- Tackling Copyright Issues in AI Image Generation Through Originality Estimation and Genericization
- Democratizing Chatbot Debugging: A Computational Framework for Evaluating and Explaining Inappropriate Chatbot Responses
- Brain-inspired Computing Based on Deep Learning for Human-computer Interaction: A Review
- AI That Helps Us Help Each Other: A Proactive System for Scaffolding Mentor-Novice Collaboration in Entrepreneurship Coaching
- DesignRepair: Dual-Stream Design Guideline-Aware Frontend Repair with Large Language Models
- Deep Natural Language Feature Learning for Interpretable Prediction
- LLM-ProS: Analyzing Large Language Models' Performance in Competitive Problem Solving
- Should AI Mimic People? Understanding AI-Supported Writing Technology Among Black Users
- On the attribution of confidence to large language models
- Assessing Logical Puzzle Solving in Large Language Models: Insights from a Minesweeper Case Study
- Revisiting Instruction Fine-tuned Model Evaluation to Guide Industrial Applications
- AI with Alien Content and Alien Metasemantics
- Explicitly Representing Syntax Improves Sentence-to-layout Prediction of Unexpected Situations
- HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs
- Visual Analysis of LLM-based Entity Resolution from Scientific Papers
- ForeSeer: Product Aspect Forecasting Using Temporal Graph Embedding
- AI-in-the-loop: The future of biomedical visual analytics applications in the era of AI
- Using Artificial Populations to Study Psychological Phenomena in Neural Models
- Assessing the nature of large language models: A caution against anthropocentrism
- Lil-Bevo: Explorations of Strategies for Training Language Models in More Humanlike Ways
- On the Role of Domain Experts in Creating Effective Tutoring Systems
- Large Language Models are biased to overestimate profoundness
- Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- It Ain't That Bad: Understanding the Mysterious Performance Drop in OOD Generalization for Generative Transformer Models
- Can we Trust Chatbots for now? Accuracy, reproducibility, traceability; a Case Study on Leonardo da Vinci's Contribution to Astronomy
- Linguacodus: A Synergistic Framework for Transformative Code Generation in Machine Learning Pipelines
- Can "consciousness" be observed from large language model (LLM) internal states? Dissecting LLM representations obtained from Theory of Mind test with Integrated Information Theory and Span Representation analysis