Evaluating Large Language Models in Theory of Mind Tasks
arXiv:2302.02083 · doi:10.1073/pnas.2405460121
Abstract
Eleven Large Language Models (LLMs) were assessed using a custom-made battery of false-belief tasks, considered a gold standard in testing Theory of Mind (ToM) in humans. The battery included 640 prompts spread across 40 diverse tasks, each one including a false-belief scenario, three closely matched true-belief control scenarios, and the reversed versions of all four. To solve a single task, a model needed to correctly answer 16 prompts across all eight scenarios. Smaller and older models solved no tasks; GPT-3-davinci-003 (from November 2022) and ChatGPT-3.5-turbo (from March 2023) solved 20% of the tasks; ChatGPT-4 (from June 2023) solved 75% of the tasks, matching the performance of six-year-old children observed in past studies. We explore the potential interpretation of these findings, including the intriguing possibility that ToM, previously considered exclusive to humans, may have spontaneously emerged as a byproduct of LLMs' improving language skills.
TRY RUNNING ToM EXPERIMENTS ON YOUR OWN: The code and tasks used in this study are available at Colab (https://colab.research.google.com/drive/1ZRtmw87CdA4xp24DNS_Ik_uA2ypaRnoU). Don't worry if you are not an expert coder, you should be able to run this code with no-to-minimum Python skills. Or copy-paste the tasks to ChatGPT's web interface. Proceedings of the National Academy of Sciences (PNAS) 2024
References in corpus (4)
Cited by in corpus (13)
- Evaluating Large Language Models in Theory of Mind Tasks
- On the Creativity of Large Language Models
- Playing repeated games with Large Language Models
- Exploring the Frontiers of LLMs in Psychological Applications: A Comprehensive Review
- The Last JITAI? Exploring Large Language Models for Issuing Just-in-Time Adaptive Interventions: Fostering Physical Activity in a Conceptual Cardiac Rehabilitation Setting
- Chatting with Bots: AI, Speech Acts, and the Edge of Assertion
- A validity-guided workflow for robust large language model research in psychology
- Can structural correspondences ground real world representational content in Large Language Models?
- From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology
- Evaluating Theory of Mind and Internal Beliefs in LLM-Based Multi-Agent Systems
- Architectures of Error: A Philosophical Inquiry into AI and Human Code Generation
- VirtLab: An AI-Powered System for Flexible, Customizable, and Large-scale Team Simulations
- Can "consciousness" be observed from large language model (LLM) internal states? Dissecting LLM representations obtained from Theory of Mind test with Integrated Information Theory and Span Representation analysis